A method and system for standardizing public bidding information based on database
By extracting standard element words and tag matching from standard bidding texts, crawling and preprocessing target bidding information, performing grammatical structure comparison and keyword extraction, the problems of insufficient classification identification and low information extraction efficiency in standardized processing of public bidding information are solved, and efficient and accurate acquisition of bidding information is achieved.
Patent Information
- Application Number
- CN202411107496.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-07
- Filing Date
- 2024-08-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-08-13
AI Technical Summary
In the standardization processing of public bidding information, the existing technology has problems such as insufficient classification identification of labeled objects, insufficient standard improvement, and low information extraction efficiency and accuracy, which cannot meet the needs of efficient and accurate acquisition.
By extracting standard element words from standard bidding text to match the label, a database is established; target bidding information is crawled and preprocessed, and converted into bidding text; grammatical structure comparison between the bidding text and the standard bidding text, keywords are extracted and standardized.
Automatic and standardized conversion of bidding documents is realized, the efficiency of users to obtain core content of target bidding information is improved, and the extraction of efficient core words is achieved to the greatest extent, avoiding the extraction of invalid or non-important core words, and improving the subsequent training effect.
Smart Images

Figure CN119129527B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and more specifically to a method and system for standardizing public bidding information based on a database. Background Art
[0002] With the development of the current information age, information has become more and more transparent and open, and it is becoming more and more convenient for users to obtain information, especially public bidding information. As competition becomes more and more fierce, whoever can quickly and accurately obtain bidding information at the first time will be able to gain the upper hand. Therefore, it is necessary to standardize public bidding information to improve the speed at which users obtain bidding information. For example: Chinese patent CN116433179A discloses a database-based public bidding information standardization method, including the following steps: S1: confirm the nature of the bidding project; S2: prequalification; S3: scheme design; S4: bidding document review; S5: start bidding; S6: confirm the result and notify the result; S7: competitive negotiation; S8: material procurement. By confirming the nature of the bidding project, it is possible to confirm whether the nature of the bidding document is an engineering construction project, government procurement or state-owned enterprise procurement, which is convenient for subsequent targeted bidding; through prequalification, the bidders can be preliminarily screened to improve efficiency; through scheme design, the technical, economic and management characteristics of the bidding project, as well as the function, scale, quality, price, progress, service and other demand targets of the bidding project can be analyzed and mastered; by confirming the results and notifying the results, the bidding unit can confirm and notify the winning bidder within the specified time and place. US Patent US10346444B1, the invention maps an organization-specific object to a standardized organizational object representing a hypothetical version of one or more actual organizational objects by matching a first parameter with a template of a standardized object; receiving an information request from a second user of a hosted computer service, the request including a second parameter; using the second parameter to map the request to a standardized object; using the first parameter and the stored information about the standardized object, providing information about the standardized object for the second user to view. Both of the above patents solve the problem of information standardization, but there is no classification and identification of annotated objects, no improvement based on standards, and the corresponding information extraction method has defects in efficiency and accuracy, which cannot meet the needs of efficient and accurate acquisition. Summary of the invention
[0003] In order to better solve the above problems, the present invention provides a method for standardizing public bidding information based on a database, the method comprising the following steps:
[0004] Step S1: extracting N standard element words from the standard bidding text, matching the N standard element words with N tags, obtaining the corresponding relationship between the N standard element words and the N tags, and creating a database according to the corresponding relationship;
[0005] Step S2: crawling multiple target tender information from the network based on the crawling conditions by the acquisition unit, performing data processing on the target tender information through preprocessing to convert the target tender information into a tender text, comparing the grammatical structure of the tender text with that of the standard tender text, and converting the tender text into a target standard tender text according to the comparison result;
[0006] The step of performing data processing on the target bidding information through preprocessing to convert the target bidding information into a bidding text includes:
[0007] Step S21: using periods as dividing points, converting the target bidding information into multiple sentences;
[0008] Step S22: taking sentences as sub-units, extracting core words from the sentences to obtain first-category core words and second-category core words;
[0009] Step S23: using a core word expansion algorithm to expand the second type of core words to obtain third core words;
[0010] Step S24: performing the following operations on the sentence: modifying the first-category core word into the first symbol while retaining the second-category core word to form a first sentence, modifying the first-category core word into the first symbol while modifying the second core word into the third core word to form a second sentence;
[0011] Step S25: predicting a simplified word of the first symbol using a prediction model based on the contents of the first sentence and the second sentence;
[0012] Step S26: using simplified words to replace the first symbol in the first sentence and the second sentence to form a third sentence and a fourth sentence;
[0013] Step S27: All the third statements and the fourth statements are combined to form the bidding text.
[0014] Step S3: extracting N keywords corresponding to the N tags from the target standard bidding text based on the positions of the N standard element words in the standard bidding text, acquiring a first standard keyword corresponding to the first keyword based on a first keyword among the N keywords and a publishing address of the target bidding information, acquiring a second standard keyword corresponding to the second keyword based on the industry field where the user is located and according to a second keyword among the N keywords and an approximate word model corresponding to the industry field; and acquiring a third standard keyword corresponding to N-2 keywords other than the first keyword and the second keyword according to the industry standard corresponding to the industry field;
[0015] Among them, the standard keywords corresponding to the N keywords include the first standard keyword, the second standard keyword and the third standard keyword.
[0016] As a more preferred technical solution of the present invention, the N standard element words include the core content of the standard bidding document, the N tags are the categories of the N standard element words, the first keyword is the name of the bidding unit, and the second keyword is the bidding product name and the bidding product ingredient name in the target bidding information.
[0017] As a more preferred technical solution of the present invention, the step S3 further includes a step S4:
[0018] Based on the first standard keyword, the second standard keyword and the third standard keyword, the N tags are associated and a first node is created, the first node is saved in the first area of the database, all the first nodes in the first area are sorted, and displayed on a map according to the timeline and the first standard keyword, the first keyword is searched in the second area of the database, and the first keyword and the second keyword and the association relationship between the first standard keyword and the second standard keyword are saved in the second area of the database according to the search result.
[0019] As a more preferred technical solution of the present invention, step S2 includes the following steps:
[0020] The data processing of the target bidding information through preprocessing to convert the target bidding information into a bidding text also includes:
[0021] Step S28: Acquire the first-category core words and the second-category core words of all sentences, form a first core word library based on all the first-category core words, and form a second core word library based on all the second-category core words;
[0022] Step S29: Count the frequency of occurrence of each first core word in the first core word library, sort the frequency in descending order, and select the N first core words with the highest frequency as candidate core words;
[0023] Step S30: Randomly combine each candidate core word with all the second-category core words in the second core word library, identify the semantic similarity between the combination and the third sentence and / or the fourth sentence, and set the corresponding candidate core word as the modified core word when the semantic similarity is greater than the first threshold.
[0024] Step S31: using the modified core words to replace some simplified words in the third sentence and the fourth sentence;
[0025] Step S32: Perform grammatical analysis on the tender text to obtain the grammatical structure and grammatical components of the tender text, and compare the grammatical structure and grammatical components with the standard grammatical structure and standard grammatical components of the standard tender text, and adjust the positions of the grammatical components of the tender text based on the comparison results and the positions of the standard grammatical components in the standard tender text to obtain the target standard tender text corresponding to the tender text.
[0026] As a more preferred technical solution of the present invention, step S3 includes the following steps:
[0027] Step S311: extracting the N keywords from the target standard tender text based on the positions of the N standard element words in the standard tender text, wherein the positions are the positions of each of the N standard element words in the standard tender text divided according to standard grammatical components;
[0028] Step S312: acquiring the homepage information of the tendering unit corresponding to the target tendering information based on the publishing address of the target tendering information, and crawling the first standard keyword of the first keyword from the homepage information;
[0029] Step S313: inputting the second keyword into the approximate word model of the industry field to which the user belongs to obtain the second standard keyword corresponding to the second keyword, wherein the approximate word model is based on the industry field corresponding to the bidding information;
[0030] Step S314: the industry standard is also obtained through the industry field, and the other N-2 keywords of the N keywords except the first keyword and the second keyword are standardized based on the industry standard, and the third standard keyword is obtained.
[0031] As a more preferred technical solution of the present invention, the step S3 also includes: after obtaining the N keywords and before obtaining the first standard keyword, searching for the first keyword in the second area of the database; when the first keyword is found, directly obtaining the first standard keyword through the correspondence between the first keyword and the first standard keyword, searching for the second keyword in the node of the first keyword, and obtaining the second standard keyword based on the second keyword; when the second keyword cannot be found and the second standard keyword cannot be obtained through the approximate word model, the second keyword is used as the second standard keyword, and the second keyword is saved as a temporary keyword in the third area of the database; when the temporary keyword is greater than a set value, the approximate word model is retrained.
[0032] As a more preferred technical solution of the present invention, step S4 includes the following steps:
[0033] Step S41: associating the first standard keyword, the second standard keyword, and the third standard keyword with the N tags based on the corresponding relationship, creating a first node based on the first standard keyword, the second standard keyword, and the third standard keyword, and saving the first node to the first area of the database;
[0034] Step S42: all nodes in the database are also sorted according to the bidding time of the target bidding information corresponding to each node, and the address of the bidding unit represented by the first standard keyword in each node is obtained, and the contents of all the nodes are displayed on the map according to the time axis;
[0035] Step S43: Also search for the first keyword in the second area of the database. When the search result is that the first keyword is found, associate the second keyword with the first keyword, and save the second keyword and the correspondence between the second keyword and the second standard keyword to the second node in the second area where the first keyword is located. When the search result is that the first keyword cannot be found, create the second node in the second area, and save the first keyword and the correspondence between the first keyword and the first standard keyword, and the second keyword and the correspondence between the second keyword and the second standard keyword to the second node.
[0036] As a more preferred technical solution of the present invention, in step S3, obtaining the approximate word model includes:
[0037] Create the approximate word model, and obtain the product name, product ingredients, approximate words of the product name, standard name of the product name, approximate words of the product ingredients and standard name of the product ingredients in the industry field from the data dictionary according to the industry field of the user as the first training data, and also obtain the common name of the product name as the second training data, and train the approximate word model based on the first training data and the second training data
[0038] As a more preferred technical solution of the present invention, the common name is the product name that is known to professionals in the industry field but does not exist in the data dictionary.
[0039] The present invention also provides a database-based open bidding information standardization processing system, which is used to implement the above method. The system includes:
[0040] The creation unit is configured to: extract N standard element words from the standard bidding text, match the N standard element words with N tags, obtain the corresponding relationship between the N standard element words and the N tags, and create a database according to the corresponding relationship;
[0041] The acquisition unit is configured to: crawl multiple target tender information from the network based on the crawling condition through the acquisition unit, perform data processing on the target tender information through preprocessing to convert the target tender information into a tender text, compare the grammatical structure of the tender text with that of the standard tender text, and convert the tender text into a target standard tender text according to the comparison result;
[0042] The standardization unit is configured to: extract N keywords corresponding to the N tags from the target standard bidding text based on the positions of the N standard element words in the standard bidding text, obtain a first standard keyword corresponding to the first keyword based on a first keyword among the N keywords and a publishing address of the public bidding information, obtain a second standard keyword corresponding to the second keyword based on the industry field where the user is located and according to a second keyword among the N keywords and an approximate word model corresponding to the industry field; and obtain a third standard keyword corresponding to N-2 keywords other than the first keyword and the second keyword according to the industry standard corresponding to the industry field;
[0043] Among them, the standard keywords corresponding to the N keywords include the first standard keyword, the second standard keyword and the third standard keyword.
[0044] Compared with the prior art, the beneficial effects of the present invention are at least as follows:
[0045] The technical solution of the present invention can realize automatic crawling of bidding information and convert the bidding information into bidding text through the cooperation of steps S1-S4, and compare the grammatical structure of the bidding text with the standard bidding text, so as to convert it into the target standard bidding document, realize automatic and standardized conversion of the bidding document, and can cooperate with the first standard keyword, the second standard keyword and the third standard keyword to realize efficient acquisition of the core content of the target bidding information, and further cooperate with the sentences and the first and second category core words and the first symbols and simplified words to realize the extraction of efficient core words to the greatest extent, avoid the extraction of invalid or unimportant core words to reduce efficiency, and prevent the other subsequent use of the above-mentioned erroneous core words as training samples to reduce the overall learning effect, that is, when the text obtained by the above method in the present invention is used as a subsequent training sample, the training effect can be further improved, and the first and second category core word libraries and candidate core words and their combination with the second category core words are further cooperated to obtain the corrected core words, which further improves the selection effect of keywords and improves the training effect as a sample. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a flow chart of a method for standardizing public bidding information based on a database of the present invention;
[0047] Figure 2 This is a structural diagram of a database-based public bidding information standardization processing system of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0049] The present invention provides a method for standardizing public bidding information based on a database. Figure 1 As shown, the method comprises the following steps:
[0050] Step S1: extracting N standard element words from the standard bidding text, matching the N standard element words with N tags, establishing a corresponding relationship between the N standard element words and the N tags, and creating a database according to the corresponding relationship;
[0051] Specifically, since there is a large amount of public bidding information published on the Internet, and since it is an open bidding, the bidding competition pressure is very high, users need to obtain the bidding information that suits them as soon as possible, and can quickly obtain the core content of the bidding information, so that they can respond to the above bidding information as quickly as possible. By extracting the above N standard element words from the above standard bidding text, wherein the above N element words contain the necessary content of the above bidding information, such as: bidding unit, bidding products or projects, quantity of bidding products, product specifications, etc., and setting N tags according to the content of the above N standard element words, by establishing the corresponding relationship between the above N standard element words and the above N tags, the standard bidding text is converted into standard structured data, which saves the time of users extracting bidding content from bidding information and improves the efficiency of obtaining bidding content.
[0052] Step S2: crawling multiple target tender information from the network based on the crawling conditions by the acquisition unit, performing data processing on the target tender information through preprocessing to convert the target tender information into a tender text, comparing the grammatical structure of the tender text with that of the standard tender text, and converting the tender text into a target standard tender text according to the comparison result;
[0053] Specifically, since there are numerous public bidding information on the Internet, it is necessary to filter out multiple public bidding information irrelevant to the user according to the needs of the user to obtain the target bidding information, wherein the above crawling conditions are set according to the needs of the user, and the above target bidding information is converted into text information for further processing. Since the language expressions of the target bidding information released by different bidding units are different, the grammatical structure of the above bidding text is obtained through grammatical analysis, and the grammatical structure of the above bidding text is compared with the grammatical structure of the above standard bidding text. According to the above comparison result, the grammatical structure of the above bidding text is adjusted to obtain the above target standard bidding text that is consistent with the grammatical structure of the above standard bidding text, so as to lay a foundation for further extracting keywords in the above target bidding text;
[0054] Furthermore, when extracting keywords, since each sentence contains different keywords, and different keywords correspond to different information and have different corresponding importance, if keywords are directly extracted from the crawled bidding information, the efficiency and accuracy of the extraction will be reduced, resulting in errors in the information accuracy of the bidding information. Acquiring accurate bidding information is the basis for subsequent processing. For example, only when the bidding information is very consistent will the bidding documents be acquired. However, it is not easy to acquire bidding documents. On the one hand, bidding documents usually need to be purchased, and more accurate data and analysis are required after purchase to analyze whether to bid for the bidding. Inaccurate keyword acquisition may make subsequent processing invalid, increasing costs. In order to improve the efficiency and accuracy of information, the data processing of the target bidding information through pre-processing of the present invention to convert the target bidding information into a bidding text includes:
[0055] Step S21: using a period as a segmentation point, converting the target bidding information into multiple sentences; generally, a period is the end of a sentence, a text is composed of multiple paragraphs, and a paragraph may include multiple sentences, i.e., sentences, so using a period as a segmentation point can initially divide the information;
[0056] Step S22: Taking sentences as sub-units, extract core words from sentences to obtain first-category core words and second-category core words; the first-category core words may be relatively unimportant core words, while the second-category core words may be relatively important core words. For example, a sentence includes "Province A Fire Department's Township Professional Team Fire Truck Procurement Project", and core words "Province A", "Province A Fire Department", "Township", "Professional Team", "Fire Truck", "Fire Truck", etc. may be extracted. The above core words have different meanings and importance. For bidding information, bidders may pay more attention to core words such as "Province A" and "Fire Truck". Such core words can reflect key information about bidding to a certain extent, while "Province A Fire Department", "Township", and "Professional Team" can also reflect bidding information to a certain extent, but for bidders, their importance is reduced. Therefore, when extracting core words, it is necessary to extract core words such as "Province A" and "Fire Truck" as the second-category core words for judgment, while core words such as "Province A Fire Department", "Township", and "Professional Team" can be judged as the first-category core words;
[0057] Step S23: Use the core word expansion algorithm to expand the second category of core words to obtain the third core words; in actual application scenarios, although some core words have some common terms, these terms are not unique meanings, for example, the commonly used word "excavator" in this field can also be called "excavator", "earthmoving machine", "shovel", etc., so in order to expand the meaning of the bidding information, the second category of core words are expanded, and the expansion method can use the core word expansion algorithm, for example, the core word expansion software is already a well-known software and technology, which will not be described in detail here; and in order to improve the accuracy of the expansion, the core word expansion library can be customized according to the bidding company's own information, so that the expansion of the core words can be more in line with its own fields and needs;
[0058] Step S24: performing the following operations on the sentence: modifying the first-category core word into the first symbol while retaining the second-category core word to form a first sentence, modifying the first-category core word into the first symbol while modifying the second core word into the third core word to form a second sentence;
[0059] Step S25: Based on the contents of the first sentence and the second sentence, use the prediction model to predict the simplified word of the first symbol; for example, the corresponding "fire headquarters", "township", and "professional team" in "Provincial Fire Department of A Province to Township Professional Team Fire Truck Procurement Project" can be modified to "somewhere"; and when multiple identical simplified words appear consecutively, only one simplified word can be retained in subsequent processing; in order to improve the accuracy of the prediction model, the model can be trained with training data, for example, the model can use sentences including multiple core words as input data, and use sentences modified with the first symbol as output results, so as to be trained according to the training data, and the trained model can output sentences including the first symbol based on the sentences; for example, when the prediction model uses place names as learning data, such as when Province A or City B or County C or Town D is used as learning data, "somewhere" is used as the result for training, so that when these data are used as input, sentences including simplified words such as "somewhere" can be output;
[0060] Step S26: Use simplified words to replace the first symbol in the first sentence and the second sentence to form the third sentence and the fourth sentence; when selecting the actual core words, "a certain place" can be generally identified as a non-important core word, so that when selecting the subsequent core words, these words are not used as key core words, which can improve the extraction efficiency and accuracy of the subsequent core words;
[0061] S27, combining all the third statements and the fourth statements to form a bidding text; the preprocessing of the target bidding information to convert the target bidding information into the bidding text also includes:
[0062] Step S28: Obtain the first-category core words and the second-category core words of all sentences, form a first core word library based on all the first-category core words, and form a second core word library based on all the second-category core words; generally speaking, after the first-category core words and the second-category core words are classified accurately after the aforementioned steps S21-26, the extraction efficiency and accuracy can be improved, but in some cases, some core words are identified as first-category core words in some cases, which is not accurate in this case, for example, "The bidder should obtain the bidding documents at E Company, Room D, Building C, Building B, City A", this information is the address where the bidding documents can be obtained for the bidding information, if it is modified to the first symbol, it may result in the subsequent inability to obtain the address;
[0063] Step S29: Count the frequency of each first core word in the first core word library, sort the frequency in descending order, and select the N first core words with the highest frequency as candidate core words; for example, core words that appear more than 3 times can be taken as candidate core words; in order to accurately publish bidding information, the address for obtaining bidding documents is usually written in great detail, such as accurate to the house number, and will appear in multiple places such as "project overview", "obtain bidding documents", "place to submit bidding documents", etc.; in order to adapt to the bidding information field of the present invention, core words that appear more than 3 times (including 3 times) are set for extraction, so that it can be determined that even if the word is a second-category core word;
[0064] Step S30: Randomly combine each of the candidate core words with all the second-category core words in the second core word library, identify the semantic similarity between the combination and the third sentence and / or the fourth sentence, and set the candidate core word corresponding to the semantic similarity greater than the first threshold as the modified core word; however, although some core words appear multiple times, this multiple times does not reflect the importance of the core words. For example, "enterprise", "statement", "implementation opinions" and so on usually appear more than 3 times in bidding information, but these are usually not identified as second-category core words. As second-category core words, they can be combined with other second-category core words to obtain core words with obvious Sentences with clear meanings, and these sentences can have a certain similarity with the third sentence and / or the fourth sentence. For example, the core words related to the house number can form sentences with clear meanings in a sentence with "tender documents" and "obtain" determined as the second-category core words, and are similar in meaning to the sentence "obtain tender documents from a certain company in a certain place" with "a certain place". Therefore, by identifying the semantic similarity between the combination and the third sentence and / or the fourth sentence, the core words are modified; step S31: using the modified core words to replace some simplified words in the third sentence and the fourth sentence;; thereby updating the third sentence and the fourth sentence and recombining them to form a tender text.
[0065] Step S3: extracting N keywords corresponding to the N tags from the target standard bidding text based on the positions of the N standard element words in the standard bidding text, acquiring a first standard keyword corresponding to the first keyword based on a first keyword among the N keywords and a publishing address of the target bidding information, acquiring a second standard keyword corresponding to the second keyword based on the industry field where the user is located and according to a second keyword among the N keywords and an approximate word model corresponding to the industry field; and acquiring a third standard keyword corresponding to N-2 keywords other than the first keyword and the second keyword according to the industry standard corresponding to the industry field;
[0066] Among them, the standard keywords corresponding to the N keywords include the first standard keyword, the second standard keyword and the third standard keyword.
[0067] Specifically, since the grammatical structure of the above-mentioned target standard bidding text is the same as the grammatical structure of the above-mentioned standard bidding text, the grammatical component positions of the above-mentioned N keywords in the above-mentioned target standard bidding text are the same as the grammatical component positions of the above-mentioned N standard element words in the above-mentioned standard bidding text, so that N keywords corresponding to N tags are extracted from the target standard bidding text through the positions of the above-mentioned N standard element words in the standard bidding text respectively, wherein the first keyword among the above-mentioned N keywords is the name of the unit that publishes the above-mentioned target bidding information, but the unit name displayed on the target bidding information may be the abbreviation of the unit or the internal name of the unit, therefore, the homepage of the above-mentioned unit is obtained through the website that publishes the above-mentioned target bidding information, and the standard name of the above-mentioned unit is obtained from the above-mentioned homepage, that is, the first standard keyword of the above-mentioned first keyword, and the second keyword among the above-mentioned N keywords is the name of the bidding product in the target bidding information, Since the names of the same product in different units may not be the same, the product names need to be standardized, and the product names are also related to the industry field. For example, a notebook can refer to a computer or a notebook for writing. Therefore, the product names need to be standardized based on the industry field. The second standard keyword corresponding to the second keyword is obtained by obtaining the approximate word model of the industry field in which the above product is located. The N-2 keywords other than the first keyword and the second keyword among the N keywords are the model, quantity and parameters of the bidding products, which are related to the industry standards of the above bidding products. Therefore, the N-2 keywords are converted into the third standard keywords according to the industry standards. Through the above technical solution, the first standard keywords, second standard keywords and third standard keywords corresponding to the N keywords in the above target bidding information are obtained, which improves the efficiency of users in obtaining the core content of the above target bidding information.
[0068] Furthermore, the N standard element words include the core content of the standard bidding document, the N standard element words include the core content of the standard bidding document, the N tags are categories of the N standard element words, the first keyword is the name of the bidding unit, and the second keyword is the bidding product name and bidding product component name in the target bidding information.
[0069] Furthermore, the step S3 further includes a step S4:
[0070] Based on the first standard keyword, the second standard keyword and the third standard keyword, the N tags are associated and a first node is created, the first node is saved in the first area of the database, all the first nodes in the first area are sorted, and displayed on a map according to the timeline and the first standard keyword, the first keyword is searched in the second area of the database, and the first keyword and the second keyword and the association relationship between the first standard keyword and the second standard keyword are saved in the second area of the database according to the search result.
[0071] Specifically, by associating the first standard keyword, the second standard keyword and the third standard keyword corresponding to the above N keywords with N tags and creating a first node, the above target bidding information is converted into structured standard bidding information, and all the nodes in the above first area are sorted, and the addresses of the bidding units corresponding to each node are displayed on the map on the timeline, so that the user can quickly and intuitively obtain the above target bidding information, and the above first keyword, that is, the name of the bidding unit in the target bidding information, is searched in the above database. When the above first keyword can be found in the above second area, the above second keyword is associated with the first keyword. At the same time, the correspondence between the above first standard keyword and the second standard keyword and the above first keyword and the second keyword is saved in the node where the first keyword is located in the above second area, so that when the bidding information of the above bidding unit is obtained next time, the first standard keyword and the second standard keyword can be directly obtained through the first keyword in the above second area, which further improves the efficiency of users in obtaining bidding content.
[0072] Furthermore, the step S2 comprises the following steps:
[0073] Step S21: filtering multiple public bidding information according to the crawling conditions of the user through the acquisition unit to acquire multiple target bidding information;
[0074] Specifically, through the above crawling conditions set by the user, that is, the conditions for obtaining the target bidding information required by the user, the above technical solution can filter out unnecessary public bidding information and obtain the target bidding information, thereby reducing the workload of the user in extracting the target bidding information.
[0075] Step S22: The target bidding information is passed through a text conversion unit to obtain the bidding text, and a grammatical analysis is performed on the bidding text to obtain the grammatical structure and grammatical components of the bidding text, and the grammatical structure and the grammatical components are compared with the standard grammatical structure and standard grammatical components of the standard bidding text, and the positions of the grammatical components of the bidding text are adjusted based on the comparison result and the positions of the standard grammatical components in the standard bidding text to obtain the target standard bidding text corresponding to the bidding text.
[0076] Specifically, the web page where the target indicator information is located is scanned by the text conversion unit, and the target bidding information is converted into the bidding text. Since the language expression habits of the target bidding information released by different units are different, the grammatical structure of the bidding text is analyzed to obtain the grammatical structure and grammatical components of the bidding text, and the grammatical structure and the position of the grammatical components of the bidding text are adjusted based on the standard grammatical structure and the position of the standard grammatical components of the standard bidding text, so that the grammatical structure and grammatical components of the bidding text are the same as those of the standard bidding text, so as to facilitate the extraction of keywords from the bidding text, thereby laying a foundation for standardizing the target bidding information.
[0077] Furthermore, step S3 includes the following steps:
[0078] Step S311: extracting the N keywords from the target standard tender text based on the positions of the N standard element words in the standard tender text, wherein the positions are the positions of each of the N standard element words in the standard tender text divided according to standard grammatical components;
[0079] Specifically, since the grammatical structure of the above-mentioned target standard bidding text is the same as that of the above-mentioned standard bidding text, the grammatical component positions of the above-mentioned N standard element words in the above-mentioned standard bidding text are the same as the positions of the above-mentioned N keywords in the above-mentioned target standard bidding document. Therefore, through the above-mentioned technical solution, the N keywords in the above-mentioned target bidding text can be obtained, that is, the core content of the above-mentioned target bidding information.
[0080] Step S312: acquiring the homepage information of the tendering unit corresponding to the target tendering information based on the publishing address of the target tendering information, and crawling the first standard keyword of the first keyword from the homepage information;
[0081] Specifically, since the first keyword obtained from the above-mentioned target bidding text, namely the name of the bidding unit, may not be a standard name but may be an abbreviation or an internal common name, the homepage address of the bidding unit is obtained through the publishing address of the above-mentioned target bidding information, and the first standard keyword corresponding to the above-mentioned first keyword is obtained by accessing the homepage content, thereby realizing the standardization of the bidding unit name.
[0082] Step S313: inputting the second keyword into the approximate word model of the industry field to which the user belongs to obtain the second standard keyword corresponding to the second keyword, wherein the approximate word model is based on the industry field corresponding to the bidding information;
[0083] Specifically, since the same keyword has different meanings for different industries, and since the target bidding information obtained by the user is the bidding information required by the above user, the bidding product name in the above target bidding information belongs to the industry field described by the above user. Therefore, the approximate word model of the above industry field can be used to obtain the second standard keyword corresponding to the above second keyword, and the standardization of the second keyword is achieved through the above technical solution.
[0084] Step S314: the industry standard is also obtained through the industry field, and the other N-2 keywords of the N keywords except the first keyword and the second keyword are standardized based on the industry standard, and the third standard keyword is obtained.
[0085] Specifically, since the above-mentioned target bidding text includes, in addition to the name of the bidding unit, the name of the bidding product and the model name, the quantity of the bidding products, product specifications and bidding time, when compiling the relevant instructions and documentation materials of the bidding products, they must be expressed in accordance with the relevant industry standards. Therefore, the above-mentioned N-2 keywords can also be standardized according to the above-mentioned industry standards, which is also convenient for users to identify. In summary, the standardization of the above-mentioned N-2 keywords is achieved through the above-mentioned technical solution.
[0086] Furthermore, the step S3 also includes: after obtaining the N keywords and before obtaining the first standard keyword, searching for the first keyword in the second area of the database; when the first keyword is found, directly obtaining the first standard keyword through the correspondence between the first keyword and the first standard keyword, searching for the second keyword in the node of the first keyword, and obtaining the second standard keyword based on the second keyword; when the second keyword cannot be found and the second standard keyword cannot be obtained through the approximate word model, the second keyword is used as the second standard keyword, and the second keyword is saved as a temporary keyword in the third area of the database; when the temporary keyword is greater than a set value, the approximate word model is retrained.
[0087] Specifically, the first keyword and the corresponding second keyword are stored in the second area of the above-mentioned database, and the addresses of the first standard keyword and the second standard keyword corresponding to the first keyword and the second keyword are also stored in the second node where the above-mentioned first keyword is located. Therefore, the above-mentioned first standard keyword and the above-mentioned second standard keyword can be directly obtained by directly searching for the above-mentioned first keyword and the above-mentioned second keyword in the above-mentioned second area. Due to the endless emergence of new products and new product ingredients, and the lack of standard names when they first appear, when the above-mentioned new products or new product ingredients appear in the target bidding information, the standard names of the above-mentioned new products or new product ingredients cannot be obtained in the above-mentioned second area or through the approximate word model. Therefore, it is necessary to temporarily use the second keyword as the second standard keyword and save the first keyword to the third area of the above-mentioned database. After the above-mentioned new product or new product ingredient appears for a certain period of time, that is, when the above-mentioned temporarily stored keyword is greater than the set value, other new products or new product ingredients related to it will also appear. Therefore, the training data, i.e., the terminology manual, etc. are re-obtained to update the above-mentioned approximate word model, so as to obtain the standard names of the above-mentioned new products or new product ingredients and the standard names of other new products or new product ingredients related to it, thereby realizing the standardization of the above-mentioned second keyword.
[0088] Furthermore, the step S4 comprises the following steps:
[0089] Step S41: associating the first standard keyword, the second standard keyword, and the third standard keyword with the N tags based on the corresponding relationship, creating a first node based on the first standard keyword, the second standard keyword, and the third standard keyword, and saving the first node to the first area of the database;
[0090] Specifically, since the above-mentioned first standard keyword, the above-mentioned second standard keyword and the above-mentioned third standard keyword contain the core content of the above-mentioned target bidding information, the above-mentioned first standard keyword, the above-mentioned second standard keyword and the above-mentioned third standard keyword realize the standardization of the above-mentioned target bidding information, and also realize the structural standardization of the above-mentioned target bidding information by associating the above-mentioned first standard keyword, the above-mentioned second standard keyword and the above-mentioned third standard keyword with the N tags of the above-mentioned database to create a first node and save it to the first area of the above-mentioned database, which provides convenience for users to view the above-mentioned bidding information and improves the efficiency of users in obtaining target bidding information.
[0091] Step S42: all nodes in the database are also sorted according to the bidding time of the target bidding information corresponding to each node, and the address of the bidding unit represented by the first standard keyword in each node is obtained, and the contents of all the nodes are displayed on the map according to the time axis;
[0092] Specifically, since there may be multiple target bidding information and the bidding times are different, the user may miss the target bidding information that is closest in time. When there is a time conflict, the bidding units are also sorted according to their addresses. Therefore, by displaying them on the map according to the timeline, the user can obtain each target bidding information intuitively and quickly. Since the competitiveness of public bidding is relatively high, users can avoid missing the target bidding information and causing losses.
[0093] Step S43: Also search for the first keyword in the second area of the database. When the search result is that the first keyword is found, associate the second keyword with the first keyword, and save the second keyword and the correspondence between the second keyword and the second standard keyword to the second node in the second area where the first keyword is located. When the search result is that the first keyword cannot be found, create the second node in the second area, and save the first keyword and the correspondence between the first keyword and the first standard keyword, and the second keyword and the correspondence between the second keyword and the second standard keyword to the second node.
[0094] Specifically, since the target bidding information that the above-mentioned users are concerned about is all within the industry field where the above-mentioned users are located, they may pay attention to or participate in bidding for multiple products of the same above-mentioned bidding unit or multiple biddings for the same product. Therefore, in order to further improve the acquisition of structured and standardized target bidding information, the above-mentioned first keyword, i.e., the name of the bidding unit, and the second keyword, i.e., the name of the bidding product or the name of the product ingredient, are also saved in the second area of the above-mentioned database, and a second node corresponding to the first keyword is created, and the above-mentioned first keyword and the above-mentioned second keyword are associated. At the same time, the correspondence between the first standard keyword and the second standard keyword corresponding to the first keyword and the second keyword is saved in the second node. For example: the storage address of the first standard keyword and the second standard keyword is saved at the above-mentioned second node. As time goes by, a corresponding list between the first keyword and multiple second keywords will be obtained. When the target bidding information released by the same bidding unit is standardized next time, the storage location of the first marked keyword and the second standard keyword can be directly obtained by searching the above-mentioned first keyword and the second keyword in the second area of the above-mentioned database, and the first marked keyword and the second standard keyword can be obtained, thereby further improving the efficiency of users in obtaining standardized target bidding information.
[0095] Furthermore, in step S3, obtaining the approximate word model includes:
[0096] Create the approximate word model, and obtain from a data dictionary according to the industry field to which the user belongs, product names, product ingredients, approximate words of the product names, standard names of the product names, approximate words of the product ingredients, and standard names of the product ingredients in the industry field as first training data, and also obtain common names of the product names as second training data, and train the approximate word model based on the first training data and the second training data;
[0097] The commonly used names are product names that are known to professionals in the industry but do not exist in the data dictionary.
[0098] Specifically, since the same keyword has different meanings for different industries, and since the target bidding information obtained by the user is the bidding information required by the above user, the bidding products in the above target bidding information belong to the industry field to which the above user belongs. The above data dictionary can be a terminology manual for the industry field. Since there are also some bidding product names and ingredient names that are defaulted by professionals in the industry but cannot be found in the terminology manual, the above technical solution can obtain more comprehensive training data, making the above approximate word model more accurate, and thus accurately obtaining the second standard keyword corresponding to the above second keyword.
[0099] The present invention also provides a database-based open bidding information standardization processing system, the system is used to implement the above method, such as Figure 2 As shown, the system comprises:
[0100] The creation unit is configured to: extract N standard element words from the standard bidding text, match the N standard element words with N tags, obtain the corresponding relationship between the N standard element words and the N tags, and create a database according to the corresponding relationship;
[0101] The acquisition unit is configured to: crawl multiple target tender information from the network based on the crawling condition through the acquisition unit, perform data processing on the target tender information through preprocessing to convert the target tender information into a tender text, compare the grammatical structure of the tender text with that of the standard tender text, and convert the tender text into a target standard tender text according to the comparison result;
[0102] The standardization unit is configured to: extract N keywords corresponding to the N tags from the target standard bidding text based on the positions of the N standard element words in the standard bidding text, obtain a first standard keyword corresponding to the first keyword based on a first keyword among the N keywords and a publishing address of the public bidding information, obtain a second standard keyword corresponding to the second keyword based on the industry field where the user is located and according to a second keyword among the N keywords and an approximate word model corresponding to the industry field; and obtain a third standard keyword corresponding to N-2 keywords other than the first keyword and the second keyword according to the industry standard corresponding to the industry field;
[0103] Among them, the standard keywords corresponding to the N keywords include the first standard keyword, the second standard keyword and the third standard keyword.
[0104] In summary, the technical solution of the present invention can realize automatic crawling of bidding information and convert the bidding information into bidding text through the cooperation of steps S1-S4, and compare the grammatical structure of the bidding text with the standard bidding text, so as to convert it into the target standard bidding document, thereby realizing automatic and standardized conversion of the bidding document, and can cooperate with the first standard keyword, the second standard keyword and the third standard keyword to realize efficient acquisition of the core content of the target bidding information, and further cooperate with the sentences and the first and second category core words and the first symbols and simplified words to realize the extraction of efficient core words to the greatest extent, avoid invalid or unimportant core words from being extracted to reduce efficiency, and prevent the other subsequent use of the above-mentioned erroneous core words as training samples to reduce the overall learning effect, that is, when the text obtained by the above method in the present invention is used as a subsequent training sample, the training effect can be further improved, and the first and second category core word libraries and the candidate core words and their combination with the second category core words are further cooperated to obtain the corrected core words, which further improves the selection effect of the core words and improves the training effect as a sample.
[0105] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0106] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
[0107] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for standardizing public bidding information based on a database, characterized in that: The method comprises the following steps: Step S1: extracting N standard element words from the standard bidding text, matching the N standard element words with N tags, obtaining the corresponding relationship between the N standard element words and the N tags, and creating a database according to the corresponding relationship; Step S2: crawling multiple target tender information from the network based on the crawling conditions by the acquisition unit, performing data processing on the target tender information through preprocessing to convert the target tender information into a tender text, comparing the grammatical structure of the tender text with that of the standard tender text, and converting the tender text into a target standard tender text according to the comparison result; Step S3: extracting N keywords corresponding to the N tags from the target standard bidding text based on the positions of the N standard element words in the standard bidding text, acquiring a first standard keyword corresponding to the first keyword based on a first keyword among the N keywords and a publishing address of the target bidding information, acquiring a second standard keyword corresponding to the second keyword based on the industry field where the user is located and according to a second keyword among the N keywords and an approximate word model corresponding to the industry field; and acquiring a third standard keyword corresponding to N-2 keywords other than the first keyword and the second keyword according to the industry standard corresponding to the industry field; The standard keywords corresponding to the N keywords include the first standard keyword, the second standard keyword and the third standard keyword; The step of performing data processing on the target bidding information through preprocessing to convert the target bidding information into a bidding text includes: Step S21: using periods as dividing points, converting the target bidding information into multiple sentences; Step S22: taking sentences as sub-units, extracting core words from the sentences to obtain first-category core words and second-category core words; Step S23: using a core word expansion algorithm to expand the second type of core words to obtain third core words; Step S24: performing the following operations on the sentence: modifying the first-category core word into the first symbol while retaining the second-category core word to form a first sentence, modifying the first-category core word into the first symbol while modifying the second core word into the third core word to form a second sentence; Step S25: predicting a simplified word of the first symbol using a prediction model based on the contents of the first sentence and the second sentence; Step S26: using simplified words to replace the first symbol in the first sentence and the second sentence to form a third sentence and a fourth sentence; Step S27: All the third statements and the fourth statements are combined to form the bidding text.
2. The method according to claim 1, characterized in that The step S3 further includes a step S4: Based on the first standard keyword, the second standard keyword and the third standard keyword, the N tags are associated and a first node is created, the first node is saved in the first area of the database, all the first nodes in the first area are sorted, and displayed on a map according to the timeline and the first standard keyword, the first keyword is searched in the second area of the database, and the first keyword and the second keyword and the association relationship between the first standard keyword and the second standard keyword are saved in the second area of the database according to the search result.
3. The method according to claim 1, characterized in that The target bidding information is subjected to data processing by preprocessing to convert the target bidding information into the bidding text, and further includes: Step S28: Acquire the first-category core words and the second-category core words of all sentences, form a first core word library based on all the first-category core words, and form a second core word library based on all the second-category core words; Step S29: Count the frequency of occurrence of each first core word in the first core word library, sort the frequency in descending order, and select the N first core words with the highest frequency as candidate core words; Step S30: Randomly combine each candidate core word with all the second-category core words in the second core word library, identify the semantic similarity between the combination and the third sentence and / or the fourth sentence, and set the corresponding candidate core word as the modified core word when the semantic similarity is greater than the first threshold. Step S31: using the modified core words to replace some simplified words in the third sentence and the fourth sentence; Step S32: Perform grammatical analysis on the tender text to obtain the grammatical structure and grammatical components of the tender text, and compare the grammatical structure and grammatical components with the standard grammatical structure and standard grammatical components of the standard tender text, and adjust the positions of the grammatical components of the tender text based on the comparison results and the positions of the standard grammatical components in the standard tender text to obtain the target standard tender text corresponding to the tender text.
4. The method according to claim 1, characterized in that: The step S3 comprises the following steps: Step S311: extracting the N keywords from the target standard tender text based on the positions of the N standard element words in the standard tender text, wherein the positions are the positions of each of the N standard element words in the standard tender text divided according to standard grammatical components; Step S312: acquiring the homepage information of the tendering unit corresponding to the target tendering information based on the publishing address of the target tendering information, and crawling the first standard keyword of the first keyword from the homepage information; Step S313: inputting the second keyword into the approximate word model of the industry field to which the user belongs to obtain the second standard keyword corresponding to the second keyword, wherein the approximate word model is based on the industry field corresponding to the bidding information; Step S314: the industry standard is also obtained through the industry field, and the other N-2 keywords of the N keywords except the first keyword and the second keyword are standardized based on the industry standard, and the third standard keyword is obtained.
5. The method according to claim 1, characterized in that The step S3 also includes: after obtaining the N keywords and before obtaining the first standard keyword, searching for the first keyword in the second area of the database, when the first keyword is found, directly obtaining the first standard keyword through the correspondence between the first keyword and the first standard keyword, and searching for the second keyword in the node of the first keyword, and obtaining the second standard keyword based on the second keyword, when the second keyword cannot be found and the second standard keyword cannot be obtained through the approximate word model, using the second keyword as the second standard keyword, and saving the second keyword as a temporary keyword in the third area of the database, and retraining the approximate word model when the temporary keyword is greater than a set value.
6. The method according to claim 2, characterized in that The step S4 comprises the following steps: Step S41: associating the first standard keyword, the second standard keyword, and the third standard keyword with the N tags based on the corresponding relationship, creating a first node based on the first standard keyword, the second standard keyword, and the third standard keyword, and saving the first node to the first area of the database; Step S42: all nodes in the database are also sorted according to the bidding time of the target bidding information corresponding to each node, and the address of the bidding unit represented by the first standard keyword in each node is obtained, and the contents of all the nodes are displayed on the map according to the time axis; Step S43: Also search for the first keyword in the second area of the database. When the search result is that the first keyword is found, associate the second keyword with the first keyword, and save the second keyword and the correspondence between the second keyword and the second standard keyword to the second node in the second area where the first keyword is located. When the search result is that the first keyword cannot be found, create the second node in the second area, and save the first keyword and the correspondence between the first keyword and the first standard keyword, and the second keyword and the correspondence between the second keyword and the second standard keyword to the second node.
7. The method according to claim 1, characterized in that In the step S3, obtaining the approximate word model includes: Create the approximate word model, and obtain from the data dictionary according to the industry field to which the user belongs, and use the product names, product ingredients, approximate words of the product names, standard names of the product names, approximate words of the product ingredients and standard names of the product ingredients in the industry field as the first training data, and also obtain the common names of the product names as the second training data, and train the approximate word model based on the first training data and the second training data.
8. The method according to claim 7, characterized in that The common name is the product name that is known to professionals in the industry but does not exist in the data dictionary.
9. A database-based open bidding information standardization processing system, the system is used to implement the method according to any one of claims 1 to 8, characterized in that: The creation unit is configured to: extract N standard element words from the standard bidding text, match the N standard element words with N tags, obtain the corresponding relationship between the N standard element words and the N tags, and create a database according to the corresponding relationship; The acquisition unit is configured to: crawl multiple target tender information from the network based on the crawling condition through the acquisition unit, perform data processing on the target tender information through preprocessing to convert the target tender information into a tender text, compare the grammatical structure of the tender text with that of the standard tender text, and convert the tender text into a target standard tender text according to the comparison result; The standardization unit is configured to: extract N keywords corresponding to the N tags from the target standard bidding text based on the positions of the N standard element words in the standard bidding text, obtain a first standard keyword corresponding to the first keyword based on a first keyword among the N keywords and a publishing address of the public bidding information, obtain a second standard keyword corresponding to the second keyword based on the industry field where the user is located and according to a second keyword among the N keywords and an approximate word model corresponding to the industry field; and obtain a third standard keyword corresponding to N-2 keywords other than the first keyword and the second keyword according to the industry standard corresponding to the industry field; Among them, the standard keywords corresponding to the N keywords include the first standard keyword, the second standard keyword and the third standard keyword.
Citation Information
Patent Citations
Database-based public bid invitation information standardization method
CN116433179A
Management of standardized organizational data
US10346444B1
Information extraction method for bid inviting text
CN108874771A
Synonym acquisition method and device, equipment, storage medium and program product
CN117743501A