Duplicated text deletion method and device, computer equipment, storage medium and program product
By performing feature extraction and BERT model similarity calculation on candidate texts in bid documents, and identifying and deleting texts duplicate with bid documents, the problem of inefficient deletion of duplicate texts in the prior art is solved, and a more fair and transparent bidding process is achieved.
Patent Information
- Application Number
- CN202510348306.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
The existing duplicate text deletion methods are inefficient and cannot effectively identify and delete texts duplicate from bidding documents, resulting in bias in review results and opaque and unfair bidding process.
By extracting the candidate text in the bidding file feature, the first bag of words vector is obtained and compared with the text information in the bidding file. The similarity of the bag of words vector is calculated using the BERT model, and the duplicate text is determined and deleted.
It improves the efficiency of identifying and deleting duplicate texts in bid documents, ensuring the fairness of the review results and the transparency of the bidding process.
Smart Images

Figure CN120216668A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to a method, device, computer device, storage medium and program product for deleting duplicate texts. Background Art
[0002] In the bidding and tendering activities, there are two important documents, namely the tender document and the bid document. The bid document is a responsive document prepared by the bidder according to the requirements of the tender document, while the tender document is a document prepared and issued by the tenderer or the tendering agency to clarify the requirements, rules, evaluation criteria and contract terms of the tender project, etc.
[0003] If there are a large number of texts from the tender document in the content of the bid document, it may affect the reviewers when reviewing the bid document, resulting in deviations in the review results, and thus making the tendering process less transparent and fair. Therefore, it is necessary to delete the texts related to the tender document in the bid document. Currently, the existing methods for deleting duplicate texts are mainly implemented based on relatively complex algorithms, such as text matching algorithms, fingerprint recognition algorithms, and cosine similarity algorithms, etc.
[0004] However, the above methods for deleting duplicate texts have the problem of low efficiency. Summary of the Invention
[0005] Based on this, it is necessary to provide a method, device, computer device, storage medium and program product for deleting duplicate texts that can improve the efficiency of deleting duplicate texts in view of the above technical problems.
[0006] In a first aspect, the present application provides a method for deleting duplicate texts, including:
[0007] Performing feature extraction on each candidate text in the target bid document to obtain first bag-of-words vectors respectively corresponding to the candidate texts;
[0008] In the case where it is determined that the target candidate text is the text included in the tender document according to the first bag-of-words vectors respectively corresponding to the candidate texts and the text information in the tender document, deleting the target candidate text from the target bid document to obtain the final bid document.
[0009] In one embodiment, the performing feature extraction on each candidate text in the target bid document to obtain first bag-of-words vectors respectively corresponding to the candidate texts includes:
[0010] Segmenting the text of the target bid document to obtain multiple candidate texts corresponding to the target bid document;
[0011] Feature extraction is performed on each candidate text to obtain first bag-of-words vectors corresponding to the respective candidate texts.
[0012] In one embodiment, the above-mentioned feature extraction of each candidate text to obtain first bag-of-words vectors corresponding to the respective candidate texts includes:
[0013] Word segmentation processing is performed on each candidate text to obtain multiple word segmentation units corresponding to the respective candidate texts;
[0014] According to the multiple word segmentation units corresponding to the respective candidate texts, a vocabulary corresponding to the target tender document is constructed; each word segmentation unit in the vocabulary is different from each other;
[0015] According to the vocabulary, the first bag-of-words vectors corresponding to the respective candidate texts are determined.
[0016] In one embodiment, the above method further includes:
[0017] Feature extraction is performed on all the texts in the tender document to obtain second bag-of-words vectors corresponding to all the texts in the tender document;
[0018] For any one of the first bag-of-words vectors, the similarity between the first bag-of-words vector and the second bag-of-words vector is obtained;
[0019] According to the similarity and a preset similarity threshold, it is determined whether the target candidate text is a text included in the tender document.
[0020] In one embodiment, the above-mentioned obtaining the similarity between the first bag-of-words vector and the second bag-of-words vector includes:
[0021] The first bag-of-words vector and the second bag-of-words vector are input into a preset Bidirectional Encoder Representations from Transformers (BERT) model for similarity calculation to obtain the similarity between the first bag-of-words vector and the second bag-of-words vector.
[0022] In one embodiment, the above-mentioned determining whether the target candidate text is a text included in the tender document according to the similarity and a preset similarity threshold includes:
[0023] If the similarity meets the preset similarity threshold, it is determined that the target candidate text is a text included in the tender document;
[0024] If the similarity does not meet the preset similarity threshold, it is determined that the target candidate text is not a text included in the tender document.
[0025] In a second aspect, the present application further provides a device for deleting duplicate texts, including:
[0026] An extraction module for extracting features from each candidate text in the target tender document to obtain first bag-of-words vectors corresponding to the respective candidate texts;
[0027] A deletion module for determining that the target candidate text is the text included in the tender document according to the first bag-of-words vectors corresponding to the respective candidate texts and the text information in the tender document, and deleting the target candidate text from the target tender document to obtain a final tender document.
[0028] In a third aspect, the present application further provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0029] Extracting features from each candidate text in the target tender document to obtain first bag-of-words vectors corresponding to the respective candidate texts;
[0030] Determining that the target candidate text is the text included in the tender document according to the first bag-of-words vectors corresponding to the respective candidate texts and the text information in the tender document, and deleting the target candidate text from the target tender document to obtain a final tender document.
[0031] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0032] Extracting features from each candidate text in the target tender document to obtain first bag-of-words vectors corresponding to the respective candidate texts;
[0033] Determining that the target candidate text is the text included in the tender document according to the first bag-of-words vectors corresponding to the respective candidate texts and the text information in the tender document, and deleting the target candidate text from the target tender document to obtain a final tender document.
[0034] In a fifth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0035] Extracting features from each candidate text in the target tender document to obtain first bag-of-words vectors corresponding to the respective candidate texts;
[0036] Determining that the target candidate text is the text included in the tender document according to the first bag-of-words vectors corresponding to the respective candidate texts and the text information in the tender document, and deleting the target candidate text from the target tender document to obtain a final tender document.
[0037] The method, device, computer device, storage medium, and program product for deleting the above-mentioned duplicate text extract features from each candidate text in the tender document to obtain the feature vectors of each candidate text in the tender document, and determine the target candidate text based on the text information of the tender invitation document and the feature vectors of the candidate text, and delete the target candidate text from the target tender document to obtain the final tender document. Compared with the existing method of directly comparing the tender document and the tender invitation document for algorithmic duplicate checking, the present application extracts features from the content of the tender document, greatly improving the recognition efficiency of the text in the tender document that duplicates the content of the tender invitation document, thereby improving the deletion efficiency of the text in the tender document that duplicates the content of the tender invitation document. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is an application environment diagram of the method for deleting duplicate text in an embodiment;
[0040] Figure 2 It is a flowchart of the method for deleting duplicate text in an embodiment;
[0041] Figure 3 It is a flowchart of the method for deleting duplicate text in another embodiment;
[0042] Figure 4 It is a flowchart of the method for deleting duplicate text in another embodiment;
[0043] Figure 5 It is a flowchart of the method for deleting duplicate text in another embodiment;
[0044] Figure 6 It is a flowchart of the method for deleting duplicate text in another embodiment;
[0045] Figure 7 It is a flowchart of the method for deleting duplicate text in another embodiment;
[0046] Figure 8 It is a structural block diagram of the device for deleting duplicate text in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In order to make the purpose, technical solutions and advantages of this application more clear and understandable, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0048] In the bidding and tendering activities, there are two important documents, namely the tender document and the bid document. The bid document is a responsive document prepared by the bidder according to the requirements of the tender document, while the tender document is a document prepared and issued by the tenderer or the tendering agency, which is used to clarify the requirements, rules, evaluation criteria and contract terms of the tender project, etc.
[0049] If there is a large amount of text from the tender document in the content of the bid document, it may cause the reviewers to be influenced by the text of the tender document when evaluating the bid document, resulting in deviation of the evaluation results, and thus making the tendering process less transparent and fair. Therefore, it is necessary to delete the text related to the tender document in the bid document. At present, the existing methods for deleting duplicate text are mainly implemented based on relatively complex algorithms, such as text matching algorithms, fingerprint recognition algorithms and cosine similarity algorithms, etc.
[0050] However, the above methods for deleting duplicate text have the problem of low efficiency. This application aims to solve this problem.
[0051] After the background technology of the method for deleting duplicate text provided by the embodiments of this application is introduced above, below, the implementation environment involved in the method for deleting duplicate text provided by the embodiments of this application will be briefly described. The method for deleting duplicate text provided by the embodiments of this application can be applied to, for example Figure 1In the computer device shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for deleting duplicate text. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0052] Those skilled in the art can understand that Figure 1 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0053] After the application scenario of the method for deleting duplicate text provided in the embodiments of this application is introduced above, the method for deleting duplicate text described in this application will be introduced in detail below.
[0054] In one embodiment, as Figure 2 shown, a method for deleting duplicate text is provided. Taking the method applied to the Figure 1 computer device in as an example for illustration, the method includes the following steps:
[0055] S201. Extract features from each candidate text in the target tender document to obtain the first word bag vector corresponding to each candidate text.
[0056] Among them, the target tender document refers to any one of a large number of tender documents.
[0057] Among them, the candidate text refers to the text obtained after segmenting all the text in the target tender document according to the preset segmentation rules. For example, each paragraph of the text in the target tender document can be a candidate text, or each sentence of the text in the target tender document can be a candidate text.
[0058] Among them, the first bag-of-words vector refers to the vector after vectorizing the candidate text.
[0059] In this embodiment, when it is necessary to delete the content related to the tender document in the target tender document to ensure that the target tender document is not affected by the text of the tender document during the review process, thus avoiding the problems of opaque and unfair tendering process, it is necessary to first obtain the target tender document, segment the text of the target tender document according to the preset segmentation rules to obtain multiple candidate documents corresponding to the target tender document, and extract the features of each candidate text in the target tender document to obtain the first bag-of-words vector corresponding to each candidate text.
[0060] Optionally, each candidate text in the target tender document can be input into a pre-trained bag-of-words model for feature extraction to obtain the first bag-of-words vector corresponding to each candidate text. It should be noted that the pre-trained bag-of-words model can be trained according to the initial bag-of-words model, sample text, and the bag-of-words vector corresponding to the sample text.
[0061] S202. When it is determined that the target candidate text is the text included in the tender document based on the first bag-of-words vector corresponding to each candidate text and the text information in the tender document, delete the target candidate text from the target tender document to obtain the final tender document.
[0062] In this embodiment, after determining the first bag-of-words vector corresponding to each candidate text, the tender document can be obtained, and the features of all the text in the tender document can be extracted to obtain the bag-of-words vector corresponding to all the text in the tender document. Then, determine the similarity between the first bag-of-words vector corresponding to each candidate text and the bag-of-words vector corresponding to all the text in the tender document. When the similarity between the first bag-of-words vector corresponding to each candidate text and the bag-of-words vector corresponding to all the text in the tender document meets the preset conditions, determine the candidate text corresponding to the similarity as the target candidate text, and delete the target candidate text from the target tender document to obtain the target tender document after deleting the target candidate text, and use the target tender document after deleting the target candidate text as the final tender document.
[0063] The deletion of duplicate texts provided in this embodiment is achieved by extracting features from each candidate text in the tender document to obtain the feature vectors of each candidate text in the tender document, and determining the target candidate text based on the text information of the tender invitation document and the feature vectors of the candidate texts, and deleting the target candidate text from the target tender document to obtain the final tender document. Compared with the existing method of directly comparing the tender document and the tender invitation document for algorithmic duplicate checking, this application extracts features from the content of the tender document, greatly improving the recognition efficiency of the texts in the tender document that are duplicate with the content of the tender invitation document, thereby improving the deletion efficiency of the texts in the tender document that are duplicate with the content of the tender invitation document.
[0064] In one embodiment, based on Figure 2 the embodiment shown, the process of obtaining the first bag-of-words vectors corresponding to the candidate texts respectively can be described as follows. Figure 3 As shown, the above S201 "extract features from each candidate text in the target tender document to obtain the first bag-of-words vectors corresponding to each candidate text respectively" includes:
[0065] S301. Segment the text of the target tender document to obtain multiple candidate texts corresponding to the target tender document.
[0066] In this embodiment, after obtaining the target tender document, the text of the target tender document can be segmented according to a preset segmentation rule to obtain multiple candidate texts corresponding to the target tender document. For example, each paragraph of the entire text of the target tender document can be a candidate text, or each sentence of the entire text of the target tender document can be a candidate text.
[0067] S302. Extract features from each candidate text to obtain the first bag-of-words vectors corresponding to each candidate text respectively.
[0068] In this embodiment, after obtaining multiple candidate texts corresponding to the target tender document, features can be extracted from each candidate text respectively to obtain the first bag-of-words vectors corresponding to each candidate text respectively. For example, each candidate text can be input into a preset bag-of-words model for feature extraction to obtain the first bag-of-words vectors corresponding to each candidate text respectively. Among them, the preset bag-of-words model can be trained based on an initial bag-of-words model, sample texts, and the bag-of-words vectors corresponding to the sample texts.
[0069] Exemplarily, the example code of the bag-of-words model is as follows:
[0070] import numpy as np
[0071] from sklearn.feature_extraction.text import CountVectorizer
[0072] from sklearn.model_selection import train_test_split
[0073] Optionally, a method for obtaining the first bag-of-words vector corresponding to each candidate text is provided below. See Figure 4 , the above S302 "performing feature extraction on each candidate text to obtain the first bag-of-words vector corresponding to each candidate text" includes:
[0074] S401. Perform word segmentation on each candidate text to obtain multiple word segmentation units corresponding to each candidate text.
[0075] In this embodiment, after obtaining each candidate text, word segmentation can be performed on each candidate text according to a preset word segmentation rule to obtain multiple word segmentation units corresponding to each candidate text. Optionally, word segmentation can be performed on the candidate text according to word segmentation rules such as phrases, idioms, proper nouns, auxiliary words, verbs, adjectives, punctuation marks, etc. For example, if the candidate text is "I love learning English", the multiple word segmentation units of the candidate text include: "I", "love", "learning", and "English".
[0076] S402. Construct a vocabulary corresponding to the target tender document according to the multiple word segmentation units corresponding to each candidate text; each word segmentation unit in the vocabulary is different from each other.
[0077] In this embodiment, after obtaining the multiple word segmentation units corresponding to each candidate text, a vocabulary corresponding to the target tender document can be constructed according to the multiple word segmentation units corresponding to each candidate text, and each word segmentation unit in the vocabulary is different from each other. For example, the first candidate text is "I love learning English", the second candidate text is "I love nature", the multiple word segmentation units corresponding to the first candidate text include: "I", "love", "learning", and "English", and the multiple word segmentation units corresponding to the second candidate text include: "I", "love", and "nature". Then, according to the multiple word segmentation units corresponding to the first candidate text and the multiple word segmentation units corresponding to the second candidate text, the constructed vocabulary includes: "I", "love", "learning", "English", and "nature".
[0078] S403. Determine the first bag-of-words vector corresponding to each candidate text according to the vocabulary.
[0079] In this embodiment, after obtaining the vocabulary corresponding to the target tender document, the vector processing can be performed on each candidate text according to the vocabulary to obtain the first bag-of-words vector corresponding to each candidate text. Continuing with the above example, the vocabulary includes: "I", "love", "study", "English", and "nature". The multiple word segmentation units corresponding to the first candidate text include: "I", "love", "study", and "English". The multiple word segmentation units corresponding to the second candidate text include: "I", "love", and "nature". Then, the first bag-of-words vector corresponding to the first candidate text is [1, 1, 1, 1, 0], and the first bag-of-words vector corresponding to the second candidate text is [1, 1, 0, 0, 1].
[0080] Exemplarily, the code for obtaining the first bag-of-words vector corresponding to the candidate text is shown below:
[0081] # Example candidate text
[0082] documents =
[0083] "I like programming and machine learning.",
[0084] "Programming is a very interesting skill.",
[0085] "Machine learning is a branch of artificial intelligence.",
[0086] "I like reading and watching movies."
[0088] # Create a CountVectorizer object
[0089] vectorizer = CountVectorizer(stop_words='chinese') # Use Chinese stop words
[0090] # Convert the text data into a term frequency matrix
[0091] X = vectorizer.fit_transform(documents)
[0092] # Obtain the vocabulary
[0093] feature_names = vectorizer.get_feature_names_out()
[0094] # Print the bag-of-words vector and the vocabulary
[0095] print("Bag-of-words vector:\n", X.toarray())
[0096] print("Glossary:\n", feature_names)
[0097] The method for determining the first bag-of-words vector provided in this embodiment obtains the bag-of-words vectors corresponding to each candidate text in the tender document by sequentially performing segmentation processing and word segmentation processing on the tender document and performing vector processing on multiple word segmentation units after the word segmentation processing, providing a data basis for subsequent identification of the content related to the tender document in the tender document based on the bag-of-words vectors corresponding to the candidate texts.
[0098] In one embodiment, based on Figures 2 - 4 the embodiment shown, as Figure 5 shown, the above method further includes:
[0099] S203. Extract features from all the texts in the tender document to obtain the second bag-of-words vectors corresponding to all the texts in the tender document.
[0100] In this embodiment, after extracting features from the texts of the tender document to obtain the first bag-of-words vectors corresponding to each candidate text in the tender document, it is also necessary to extract features from all the texts in the tender document to obtain the second bag-of-words vectors corresponding to all the texts in the tender document. Optionally, all the texts of the tender document can be input into a preset bag-of-words model for feature extraction to obtain the second bag-of-words vectors corresponding to all the texts in the tender document. It should be noted that the preset bag-of-words model can be trained based on the initial bag-of-words model, sample texts, and the bag-of-words vectors corresponding to the sample texts.
[0101] S204. For any one of the first bag-of-words vectors, obtain the similarity between the first bag-of-words vector and the second bag-of-words vector.
[0102] In this embodiment, after obtaining the first bag-of-words vectors corresponding to each candidate text in the tender document and the second bag-of-words vectors corresponding to all the texts in the tender document, the similarity between the first bag-of-words vector and the second bag-of-words vector can be calculated to obtain the similarity between the first bag-of-words vector and the second bag-of-words vector.
[0103] Optionally, the process of obtaining the similarity between the first bag-of-words vector and the second bag-of-words vector can be described, that is, the above S204 "obtain the similarity between the first bag-of-words vector and the second bag-of-words vector" includes:
[0104] Input the first bag-of-words vector and the second bag-of-words vector into a preset Bidirectional Encoder Representations from Transformers (BERT) model for similarity calculation to obtain the similarity between the first bag-of-words vector and the second bag-of-words vector.
[0105] Among them, the preset BERT model can calculate the similarity between the input first bag-of-words vector and the second bag-of-words vector, and obtain the similarity between the first bag-of-words vector and the second bag-of-words vector.
[0106] In this embodiment, after obtaining the first bag-of-words vectors corresponding to the candidate texts in the tender documents and the second bag-of-words vectors corresponding to all the texts in the bidding documents, the first bag-of-words vectors and the second bag-of-words vectors can be respectively input into the BERT model for similarity calculation to obtain the similarity between each first bag-of-words vector and the second bag-of-words vector.
[0107] Exemplarily, the code for training the preset BERT model is shown as follows:
[0108] import torch
[0109] from transformers import BertTokenizer, BertForSequenceClassification, Trainer, TrainingArguments
[0110] from sklearn.model_selection import train_test_split
[0111] from sklearn.metrics import accuracy_score, precision_recall_fscore_support
[0112] import pandas as pd
[0113] # Assume the data has been loaded into a DataFrame, with two columns: 'text' and 'label'
[0114] # 'text' is the paragraph in the tender document, 'label' is 0 (not the content of the bidding document) or 1 (the content of the bidding document)
[0115] data = pd.read_csv('your_data.csv')
[0116] # Split the data into training set and test set
[0117] train_texts, val_texts, train_labels, val_labels = train_test_split(data['text'], data['label'], test_size=0.2, random_state=42)
[0118] # Initialize the BERT tokenizer
[0119] tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
[0120] # Tokenize the input data
[0121] def tokenize_data(texts, labels, max_length=128):
[0122] tokenized_inputs = tokenizer(list(texts), truncation=True, padding=True, max_length=max_length, return_tensors="pt")
[0123] labels = torch.tensor(labels)
[0124] return tokenized_inputs, labels
[0125] train_inputs, train_labels = tokenize_data(train_texts, train_labels)
[0126] val_inputs, val_labels = tokenize_data(val_texts, val_labels)
[0127] # Load the pre-trained BERT model for text classification
[0128] model = BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2)
[0129] # Define training parameters
[0130] training_args = TrainingArguments(
[0131] output_dir='. / results',
[0132] num_train_epochs=3,
[0133] per_device_train_batch_size=8,
[0134] per_device_eval_batch_size=8,
[0135] warmup_steps=500,
[0136] weight_decay=0.01,
[0137] logging_dir='. / logs',
[0138] logging_steps=10,
[0139] evaluation_strategy="epoch" )
[0141] # Define Trainer
[0142] trainer = Trainer(
[0143] model=model,
[0144] args=training_args,
[0145] train_dataset={'input_ids': train_inputs['input_ids'], 'attention_mask': train_inputs['attention_mask'], 'labels': train_labels},
[0146] eval_dataset={'input_ids': val_inputs['input_ids'], 'attention_mask': val_inputs['attention_mask'], 'labels': val_labels} )
[0148] # Train the model
[0149] trainer.train()
[0150] # Evaluate the model
[0151] eval_results = trainer.evaluate()
[0152] print(f"Evaluation results: {eval_results}")
[0153] # Make predictions using the model
[0154] def predict(texts):
[0155] tokenized_inputs = tokenizer(list(texts), truncation=True,padding=True, max_length=128, return_tensors="pt")
[0156] outputs = model(**tokenized_inputs)
[0157] logits = outputs.logits
[0158] preds = torch.argmax(logits, dim=-1)
[0159] return preds
[0160] Exemplarily, the code for determining the similarity between the first bag-of-words vector and the second bag-of-words vector using the trained BERT model is shown below:
[0161] First, use the Transformers library on Hugging Face to extract the embeddings of each input sentence.
[0162] The example code is as follows:
[0163] from transformers import BertTokenizer, BertModel
[0164] import torch
[0165] # Load the pre-trained model
[0166] model = BertModel.from_pretrained('bert-base-chinese', output_hidden_states=True)
[0167] tokenizer = BertTokenizer.from_pretrained('bert-base-chinese')
[0168] # Input text
[0169] sentence = 'The weather is really nice today'
[0170] # Convert the text into tokens using the bag-of-words model
[0171] tokens = tokenizer.tokenize(sentence)
[0172] print('tokens:', tokens)
[0173] # Add special symbols: [CLS] and [SEP]
[0174] tokens = ['[CLS]'] + tokens + ['[SEP]']
[0175] # Assume the maximum length we need is 10. If the length of tokens is less than 10, pad it with [PAD]
[0176] if len(tokens) < 10:
[0177] tokens = tokens + ['[PAD]' for _ in range(10 - len(tokens))]
[0178] print('tokens:', tokens)
[0179] # Create the attention mask vector
[0180] attention_mask = [1 if token != '[PAD]' else 0 for token in tokens]
[0181] # Convert all tokens to token IDs
[0182] tokens_id = tokenizer.convert_tokens_to_ids(tokens)
[0183] print('tokens_id:',tokens_id)
[0184] # Convert the token IDs and attention mask to tensors
[0185] tokens_id=torch.tensor(tokens_id).unsqueeze(0)
[0186] attention_mask=torch.tensor(attention_mask).unsqueeze(0)
[0187] # Obtain the BERT output of the text
[0188] output=model(tokens_id,attention_mask)
[0189] print(output.last_hidden_state.shape)
[0190] print(output.hidden_states.shape)
[0191] print(output.pooler_output.shape)
[0192] The output contains three types of variables, specifically as follows:
[0193] last_hidden_state: Contains the features of all tokens obtained from the last encoder;
[0194] pooler_output represents the features of the [CLS] token from the last encoder;
[0195] hidden_states contains the features of all tokens obtained from all encoder layers;
[0196] Secondly, calculate the text similarity
[0197] By using the [CLS] token output by the BERT model to represent the semantic information of the entire sentence, the text semantic similarity is calculated using the Bert vector.
[0198] The example code is as follows:
[0199] from transformers import BertTokenizer, BertModel
[0200] import torch
[0201] import numpy as np
[0202] from sklearn.metrics.pairwise import cosine_similarity
[0203] # 1. Load the BERT model and tokenizer
[0204] tokenizer = BertTokenizer.from_pretrained('bert-base-chinese')
[0205] model = BertModel.from_pretrained('bert-base-chinese')
[0206] # 2. Define a function to calculate text similarity
[0207] def get_sentence_embedding(sentence):
[0208] # Encode the input sentence
[0209] inputs = tokenizer(sentence, return_tensors='pt', max_length=128,
[0210] truncation=True, padding='max_length')
[0211] # Use the BERT model to get the output
[0212] with torch.no_grad():
[0213] outputs = model(**inputs)
[0214] # Extract the vector of the [CLS] token (sentence-level vector)
[0215] cls_embedding = outputs.last_hidden_state[:, 0, :].numpy()
[0216] return cls_embedding
[0217] def calculate_similarity(text1, text2):
[0218] # Calculate the embedding vectors of two texts
[0219] embedding1 = get_sentence_embedding(text1)
[0220] embedding2 = get_sentence_embedding(text2)
[0221] # Calculate cosine similarity
[0222] similarity = cosine_similarity(embedding1, embedding2)
[0223] return similarity[0][0]
[0224] # 3. Example texts
[0225] text1 = "The weather is really nice today"
[0226] text2 = "The weather today is extremely good"
[0227] # 4. Calculate similarity
[0228] similarity_score = calculate_similarity(text1, text2)
[0229] print(f"Similarity: {similarity_score:.4f}")
[0230] S205. Determine whether the target candidate text is included in the tender document according to the similarity and the preset similarity threshold.
[0231] Among them, the preset similarity threshold refers to the threshold preset for evaluating whether the tender document and the bidding document are consistent. It should be noted that if the similarity exceeds the preset similarity threshold, it is determined that the tender document and the bidding document are consistent; if the similarity does not exceed the preset similarity threshold, it is determined that the tender document and the bidding document are inconsistent.
[0232] Exemplarily, the code implementation for identifying similar texts in the tender document and the bidding document through the BERT model is as follows:
[0233] # Load the pre-trained BERT model for text classification
[0234] model = BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2)
[0235] # Use the model for prediction (new bidding documents need to be processed)
[0236] def predict_and_filter(texts, model, tokenizer, max_length=512):
[0237] tokenized_inputs = tokenizer(texts, truncation=True, padding=True, max_length=max_length, return_tensors="pt")
[0238] with torch.no_grad():
[0239] outputs = model(**tokenized_inputs)
[0240] logits = outputs.logits
[0241] preds = torch.argmax(logits, dim=-1)
[0242] # Assume 0 represents non-bidding document content and 1 represents bidding document content
[0243] filtered_texts = [text if pred.item() == 0 else "" for text, pred in zip(texts, preds)]
[0244] return filtered_texts
[0245] # New bidding documents need to be processed
[0246] new_bidding_file_texts = ["Here is the text content of the bidding document, which may contain parts of the bidding document..."]
[0247] filtered_texts = predict_and_filter(new_bidding_file_texts, model,tokenizer)
[0248] print(filtered_texts)
[0249] In this embodiment, after obtaining the similarity between each first bag-of-words vector and the second bag-of-words vector, the similarity can be compared with a preset similarity threshold to obtain a comparison result, and based on the comparison result, it is determined whether the target candidate text is the text included in the bidding document; optionally, if the comparison result is that the similarity is greater than the preset similarity threshold, it is determined that the target candidate text is the text in the bidding document, and if the comparison result is that the similarity is not greater than the preset similarity threshold, it is determined that the target candidate text is not the text in the bidding document.
[0250] Optionally, the following provides a method for determining whether the target candidate text is the text included in the bidding document according to the similarity and the preset similarity threshold. See Figure 6 That is, the above S205 "determine whether the target candidate text is the text included in the bidding document according to the similarity and the preset similarity threshold" includes:
[0251] S501. If the similarity meets the preset similarity threshold, it is determined that the target candidate text is the text included in the bidding document.
[0252] Among them, the preset similarity threshold can be a specific value or a numerical range.
[0253] In this embodiment, after determining the similarity between the first bag-of-words vector and the second bag-of-words vector, it is judged whether the similarity meets the preset similarity threshold. If the similarity meets the preset similarity threshold, it is determined that the target candidate text is the text included in the bidding document.
[0254] S502. If the similarity does not meet the preset similarity threshold, it is determined that the target candidate text is not the text included in the bidding document.
[0255] In this embodiment, after determining the similarity between the first bag-of-words vector and the second bag-of-words vector, it is judged whether the similarity meets the preset similarity threshold. If the similarity does not meet the preset similarity threshold, it is determined that the target candidate text is not the text included in the bidding document.
[0256] The method provided in this embodiment for determining whether the target candidate text is the text included in the bidding document provides a data basis for subsequently deleting the candidate text from the target bidding document, so that when reviewing the target bidding document, it is not affected by the bidding document, ensuring the review accuracy of the target bidding document.
[0257] In one embodiment, referring to Figure 7 , a method for deleting duplicate text is further provided, including:
[0258] S10. Segment the text of the target tender document to obtain a plurality of candidate texts corresponding to the target tender document;
[0259] S11. Perform word segmentation on each candidate text to obtain a plurality of word segmentation units corresponding to each candidate text respectively;
[0260] S12. Construct a vocabulary corresponding to the target tender document according to the plurality of word segmentation units corresponding to each candidate text respectively; each word segmentation unit in the vocabulary is different from each other;
[0261] S13. Determine the first bag-of-words vector corresponding to each candidate text according to the vocabulary;
[0262] S14. Extract features from all the text in the tender invitation document to obtain a second bag-of-words vector corresponding to all the text in the tender invitation document;
[0263] S15. For any one of the first bag-of-words vectors, input the first bag-of-words vector and the second bag-of-words vector into a pre-set Bidirectional Encoder Representations from Transformers (BERT) model for similarity calculation to obtain the similarity between the first bag-of-words vector and the second bag-of-words vector;
[0264] S16. If the similarity meets a pre-set similarity threshold, determine that the target candidate text is the text included in the tender invitation document;
[0265] S17. If the similarity does not meet the pre-set similarity threshold, determine that the target candidate text is not the text included in the tender invitation document;
[0266] S18. In the case of determining that the target candidate text is the text included in the tender invitation document, delete the target candidate text from the target tender document to obtain the final tender document.
[0267] In the duplicate text deletion provided in this embodiment, by extracting features from each candidate text in the tender document, feature vectors of each candidate text in the tender document are obtained, and based on the text information of the tender invitation document and the feature vectors of the candidate texts, the target candidate text is determined, and the target candidate text is deleted from the target tender document to obtain the final tender document. Compared with the existing method of directly comparing the tender document and the tender invitation document for algorithmic duplicate checking, in this application, by extracting features from the content of the tender document, the recognition efficiency of the text in the tender document that is repeated with the content of the tender invitation document is greatly improved, thereby improving the deletion efficiency of the text in the tender document that is repeated with the content of the tender invitation document.
[0268] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0269] Based on the same inventive concept, an embodiment of the present application also provides a duplicate text deletion device for implementing the duplicate text deletion method described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the duplicate text deletion device provided below can refer to the limitations on the duplicate text deletion method in the above text, and will not be repeated here.
[0270] In an exemplary embodiment, as Figure 8 shown, a duplicate text deletion device is provided, including: an extraction module 10 and a deletion module 11, where:
[0271] The extraction module 10 is used to extract features from each candidate text in the target tender document to obtain the first bag-of-words vector corresponding to each candidate text.
[0272] The deletion module 11 is used to determine that the target candidate text is the text included in the tender document according to the first bag-of-words vector corresponding to each candidate text and the text information in the tender document, and delete the target candidate text from the target tender document to obtain the final tender document.
[0273] In an exemplary embodiment, the extraction module 10 includes: a processing unit and an extraction unit, where:
[0274] The processing unit is specifically used to segment the text of the target tender document to obtain multiple candidate texts corresponding to the target tender document;
[0275] The extraction unit is specifically used to extract features from each candidate text to obtain the first bag-of-words vector corresponding to each candidate text.
[0276] In an exemplary embodiment, the above-mentioned extraction unit is specifically configured to perform word segmentation on each candidate text to obtain a plurality of word segmentation units respectively corresponding to each candidate text; construct a vocabulary corresponding to the target tender document according to the plurality of word segmentation units respectively corresponding to each candidate text; each word segmentation unit in the vocabulary is different from each other; determine the first bag-of-words vector corresponding to each candidate text according to the vocabulary.
[0277] In an exemplary embodiment, the above-mentioned device further includes: a feature extraction module, an acquisition module, and a determination module, where:
[0278] The feature extraction module is configured to perform feature extraction on all texts in the tender document to obtain a second bag-of-words vector corresponding to all texts in the tender document;
[0279] The acquisition module is configured to, for any one of the first bag-of-words vectors, acquire the similarity between the first bag-of-words vector and the second bag-of-words vector;
[0280] The determination module is configured to determine whether the target candidate text is the text included in the tender document according to the similarity and a preset similarity threshold.
[0281] In an exemplary embodiment, the above-mentioned acquisition module is further configured to input the first bag-of-words vector and the second bag-of-words vector into a preset Bidirectional Encoder Representations from Transformers (BERT) model for similarity calculation to obtain the similarity between the first bag-of-words vector and the second bag-of-words vector.
[0282] In an exemplary embodiment, the determination module includes: a first determination unit and a second determination unit, where:
[0283] The first determination unit is specifically configured to determine that the target candidate text is the text included in the tender document when the similarity meets the preset similarity threshold;
[0284] The second determination unit is specifically configured to determine that the target candidate text is not the text included in the tender document when the similarity does not meet the preset similarity threshold.
[0285] Each module in the above-mentioned duplicate text deletion device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0286] In an exemplary embodiment, a computer device is provided, including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0287] Extract features from each candidate text in the target tender document to obtain the first bag-of-words vectors corresponding to each candidate text respectively;
[0288] In the case where, according to the first bag-of-words vectors corresponding to each candidate text and the text information in the tender document, it is determined that the target candidate text is the text included in the tender document, delete the target candidate text from the target tender document to obtain the final tender document.
[0289] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0290] Segment the text of the target tender document to obtain multiple candidate texts corresponding to the target tender document;
[0291] Extract features from each candidate text to obtain the first bag-of-words vectors corresponding to each candidate text respectively.
[0292] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0293] Perform word segmentation on each candidate text to obtain multiple word segmentation units corresponding to each candidate text respectively;
[0294] Construct a vocabulary corresponding to the target tender document according to the multiple word segmentation units corresponding to each candidate text; each word segmentation unit in the vocabulary is different from each other;
[0295] Determine the first bag-of-words vectors corresponding to each candidate text according to the vocabulary.
[0296] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0297] Extract features from all the text in the tender document to obtain the second bag-of-words vectors corresponding to all the text in the tender document;
[0298] For any one of the first bag-of-words vectors, obtain the similarity between the first bag-of-words vector and the second bag-of-words vector;
[0299] Determine whether the target candidate text is the text included in the tender document according to the similarity and the preset similarity threshold.
[0300] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0301] Input the first bag-of-words vector and the second bag-of-words vector into a preset Bidirectional Encoder Representations from Transformers (BERT) model for similarity calculation to obtain the similarity between the first bag-of-words vector and the second bag-of-words vector.
[0302] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0303] If the similarity meets the preset similarity threshold, determine that the target candidate text is the text included in the tender document;
[0304] If the similarity does not meet the preset similarity threshold, determine that the target candidate text is not the text included in the tender document.
[0305] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0306] Extract features from each candidate text in the target tender document to obtain first bag-of-words vectors respectively corresponding to the candidate texts;
[0307] In the case where it is determined that the target candidate text is the text included in the tender document according to the first bag-of-words vectors respectively corresponding to the candidate texts and the text information in the tender document, delete the target candidate text from the target tender document to obtain the final tender document.
[0308] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0309] Segment the text of the target tender document to obtain multiple candidate texts corresponding to the target tender document;
[0310] Extract features from each candidate text to obtain first bag-of-words vectors respectively corresponding to the candidate texts.
[0311] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0312] Perform word segmentation on each candidate text to obtain multiple word segmentation units respectively corresponding to the candidate texts;
[0313] Construct a vocabulary corresponding to the target tender document according to the multiple word segmentation units respectively corresponding to the candidate texts; each word segmentation unit in the vocabulary is different from each other;
[0314] Determine the first bag-of-words vectors corresponding to the candidate texts according to the vocabulary.
[0315] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0316] Extract features from all the texts in the tender document to obtain second bag-of-words vectors corresponding to all the texts in the tender document;
[0317] For any one of the first bag-of-words vectors, obtain the similarity between the first bag-of-words vector and the second bag-of-words vector;
[0318] Determine whether the target candidate text is the text included in the tender document according to the similarity and the preset similarity threshold.
[0319] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0320] Input the first bag-of-words vector and the second bag-of-words vector into a preset Bidirectional Encoder Representations from Transformers (BERT) model for similarity calculation to obtain the similarity between the first bag-of-words vector and the second bag-of-words vector.
[0321] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0322] If the similarity meets the preset similarity threshold, determine that the target candidate text is the text included in the tender document;
[0323] If the similarity does not meet the preset similarity threshold, determine that the target candidate text is not the text included in the tender document.
[0324] In one embodiment, a computer program product is provided, including a computer program, which when executed by a processor, implements the following steps:
[0325] Extract features from each candidate text in the target tender document to obtain the first bag-of-words vector corresponding to each candidate text;
[0326] In the case of determining that the target candidate text is the text included in the tender document according to the first bag-of-words vector corresponding to each candidate text and the text information in the tender document, delete the target candidate text from the target tender document to obtain the final tender document.
[0327] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0328] Segment the text of the target tender document to obtain multiple candidate texts corresponding to the target tender document;
[0329] Extract features from each candidate text to obtain the first bag-of-words vector corresponding to each candidate text.
[0330] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0331] Perform word segmentation on each candidate text to obtain multiple word segmentation units corresponding to each candidate text;
[0332] Construct a vocabulary corresponding to the target tender document according to the multiple word segmentation units corresponding to each candidate text; each word segmentation unit in the vocabulary is different from each other;
[0333] Determine the first bag-of-words vector corresponding to each candidate text according to the vocabulary list.
[0334] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0335] Extract features from all the text in the tender document to obtain the second bag-of-words vector corresponding to all the text in the tender document;
[0336] For any first bag-of-words vector, obtain the similarity between the first bag-of-words vector and the second bag-of-words vector;
[0337] According to the similarity and the preset similarity threshold, determine whether the target candidate text is the text included in the tender document.
[0338] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0339] Input the first bag-of-words vector and the second bag-of-words vector into a preset Bidirectional Encoder Representations from Transformers (BERT) model for similarity calculation to obtain the similarity between the first bag-of-words vector and the second bag-of-words vector.
[0340] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0341] If the similarity meets the preset similarity threshold, determine that the target candidate text is the text included in the tender document;
[0342] If the similarity does not meet the preset similarity threshold, determine that the target candidate text is not the text included in the tender document.
[0343] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0344] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.
[0345] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for deleting duplicate text, characterized in that: The method comprises: Perform feature extraction on each candidate text in the target bidding document to obtain a first bag-of-words vector corresponding to each candidate text; When the target candidate text is determined to be the text included in the bidding document according to the first bag-of-words vectors corresponding to each of the candidate texts and the text information in the bidding document, the target candidate text is deleted from the target bidding document to obtain a final bidding document.
2. The method according to claim 1, characterized in that The feature extraction of each candidate text in the target bidding document to obtain the first bag-of-words vector corresponding to each candidate text includes: Segmenting the text of the target bidding document to obtain a plurality of candidate texts corresponding to the target bidding document; Feature extraction is performed on each of the candidate texts to obtain a first bag-of-words vector corresponding to each of the candidate texts.
3. The method according to claim 2, characterized in that The step of extracting features from each candidate text to obtain a first bag-of-words vector corresponding to each candidate text includes: Performing word segmentation processing on each of the candidate texts to obtain a plurality of word segmentation units corresponding to each of the candidate texts; Constructing a vocabulary corresponding to the target bidding document according to the plurality of word segmentation units corresponding to each of the candidate texts; each word segmentation unit in the vocabulary is different from each other; According to the vocabulary, a first bag-of-words vector corresponding to each candidate text is determined.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Performing feature extraction on all texts in the bidding document to obtain a second bag-of-words vector corresponding to all texts in the bidding document; For any first bag-of-words vector, obtaining the similarity between the first bag-of-words vector and the second bag-of-words vector; According to the similarity and a preset similarity threshold, it is determined whether the target candidate text is a text included in the bidding document.
5. The method according to claim 4, characterized in that The obtaining the similarity between the first bag-of-words vector and the second bag-of-words vector includes: The first bag-of-words vector and the second bag-of-words vector are input into a preset bidirectional encoder transformer representation BERT model to perform similarity calculation to obtain the similarity between the first bag-of-words vector and the second bag-of-words vector.
6. The method according to claim 4, characterized in that The determining, based on the similarity and a preset similarity threshold, whether the target candidate text is a text included in the bidding document comprises: If the similarity satisfies the preset similarity threshold, determining that the target candidate text is the text included in the bidding document; If the similarity does not satisfy the preset similarity threshold, it is determined that the target candidate text is not a text included in the bidding document.
7. A device for deleting duplicate text, characterized in that: The device comprises: An extraction module, used for performing feature extraction on each candidate text in the target bidding document to obtain a first bag-of-words vector corresponding to each candidate text; A deletion module is used to determine, based on the first bag-of-words vectors corresponding to each of the candidate texts and the text information in the tender document, that the target candidate text is the text included in the tender document, and delete the target candidate text from the target tender document to obtain a final tender document.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.