A method, apparatus, device, storage medium and program product for document processing

CN122777705APending Publication Date: 2026-09-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510323193.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

然而,在某些文档所出现的高频词语之间可能具有高度重合的情况下,如果采用当前去重方案对高频词语进行局部敏感哈希计算,所计算得到的汉明距离也是极其接近的,这就容易导致汉明距离相近的文档会被误判为相似从而被执行去重处理

Benefits of technology

[0059]In this embodiment, the document to be processed is first acquired, and its document tags are determined. These document tags describe the semantic domain to which the document belongs. Additionally, the document to be processed is segmented into multiple text segments. Then, based on the document tags, a target domain thesaurus matching the document is determined from multiple candidate domain thesauruses, and the word weight of each text segment is determined from the target domain thesaurus. The word weights describe the importance of the corresponding text segment in the semantic domain. After determining the word weights of each text segment, a target hash signature is determined based on the multiple text segments and their word weights. The target hash signature of the document to be processed is then compared with the hash signatures of multiple candidate documents in the candidate document library to determine the comparison result. This comparison result reflects whether the document to be processed and the candidate documents are similar documents. Through this method, this application pre-sets multiple candidate domain thesauruses based on a large number of documents, and different candidate domain thesauruses correspond to different semantic domains. In this way, after classifying the documents to be processed according to document tags, this application can determine the word weights of each text segment in the document based on the domain lexicon of the semantic domain. Then, it performs subsequent word segmentation weighting, hash signature calculation, and comparison according to the word weights of its respective domain. Thus, even if documents have highly overlapping word segments, the calculated hash signatures will be different due to the different word weights of their semantic domains. This effectively avoids misclassifying a large number of documents under the same category as similar, reduces the probability of misclassification, significantly improves the efficiency and accuracy of comparison, and reduces the loss of document information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777705A_ABST
    Figure CN122777705A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a document processing method, device, equipment, storage medium and program product, which effectively avoid a large number of documents in the same category from being misjudged as similar, reduce the misjudgment probability, greatly improve the comparison efficiency and accuracy, and reduce the document information loss. The method comprises the following steps: obtaining a to-be-processed document, and determining a document label of the to-be-processed document; performing word segmentation on the to-be-processed document to obtain a plurality of text segments of the to-be-processed document; determining a target domain vocabulary library matched by the to-be-processed document from a plurality of candidate domain vocabulary libraries based on the document label of the to-be-processed document, and determining a word weight of each text segment of the to-be-processed document from the target domain vocabulary library; determining a target hash signature of the to-be-processed document based on the plurality of text segments of the to-be-processed document and the word weight of each text segment of the to-be-processed document; and comparing the target hash signature of the to-be-processed document with hash signatures of a plurality of candidate documents in a candidate document library to determine a comparison result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a method, apparatus, device, storage medium, and program product for document processing. Background Technology

[0002] With the rapid development of the internet and digital technologies, the number of documents has exploded. Among these documents, a large amount of content is duplicated or similar, posing a significant challenge to information retrieval, storage, and management. Deduplication of documents helps reduce the time spent filtering through large amounts of duplicate or similar documents, improves information retrieval efficiency, and also reduces storage requirements, lowers storage costs, and improves storage efficiency.

[0003] In existing document deduplication schemes, locality-sensitive hashing (simhash) is typically used to calculate the Hamming distance between the hash signatures of multiple documents, and then the Hamming distance is used to determine whether the compared documents are similar. However, in cases where high-frequency words in some documents may highly overlap, if the current deduplication scheme uses simhash to calculate the Hamming distance for these high-frequency words, the calculated distances will also be extremely close. This can easily lead to documents with similar Hamming distances being mistakenly identified as similar and thus subjected to deduplication. This not only increases the probability of document misidentification but may also result in the loss of document information due to incorrect deduplication. Summary of the Invention

[0004] This application provides a document processing method, apparatus, device, storage medium, and program product, which effectively avoids a large number of documents under the same category being misjudged as similar, reduces the probability of misjudgment, greatly improves the efficiency and accuracy of comparison, and reduces the loss of document information.

[0005] Firstly, embodiments of this application provide a document processing method. The method includes:

[0006] Obtain the document to be processed and determine the document tags of the document to be processed, wherein the document tags of the document to be processed are used to describe the semantic domain to which the document to be processed belongs;

[0007] The document to be processed is segmented into words to obtain multiple text segments of the document to be processed;

[0008] Based on the document tags of the document to be processed, the target domain thesaurus that the document to be processed matches is determined from multiple candidate domain thesauruses, and the word weight of each text segment of the document to be processed is determined from the target domain thesaurus. The word weight is used to describe the importance of the corresponding text segment in the semantic domain.

[0009] Based on multiple text segmentations of the document to be processed and the word weight of each text segment in the document to be processed, the target hash signature of the document to be processed is determined;

[0010] The target hash signature of the document to be processed is compared with the hash signatures of multiple candidate documents in the candidate document library to determine the comparison result. The comparison result is used to reflect whether the document to be processed and the candidate documents are similar documents.

[0011] Secondly, embodiments of this application provide a document processing apparatus. The document processing apparatus includes...

[0012] The acquisition unit is used to acquire the document to be processed.

[0013] A determining unit is used to determine the document tags of the document to be processed, wherein the document tags of the document to be processed are used to describe the semantic domain to which the document to be processed belongs;

[0014] The word segmentation unit is used to segment the document to be processed into multiple text words of the document to be processed.

[0015] The determining unit is configured to determine the target domain thesaurus that matches the document to be processed from multiple candidate domain thesauruses based on the document tags of the document to be processed, and to determine the word weight of each text segment of the document to be processed from the target domain thesaurus, wherein the word weight is used to describe the importance of the corresponding text segment in the semantic domain;

[0016] The determining unit is used to determine the target hash signature of the document to be processed based on multiple text segments of the document to be processed and the word weight of each text segment in the document to be processed;

[0017] The comparison unit is used to compare the target hash signature of the document to be processed with the hash signatures of multiple candidate documents in the candidate document library, and determine the comparison result. The comparison result is used to reflect whether the document to be processed and the candidate documents are similar documents.

[0018] In one possible design, in another implementation of the embodiments of this application, the comparison unit is specifically used for:

[0019] The target hash signature of the document to be processed is split into multiple target character segments arranged in order, and each target character segment includes target binary code;

[0020] The hash signature of each candidate document is split into strings to obtain multiple candidate character segments of each candidate document arranged in the order stated above, and each candidate character segment includes candidate binary code;

[0021] The comparison result is determined based on the target binary codes of multiple target character segments in the document to be processed and the candidate binary codes of multiple candidate character segments in each candidate document.

[0022] In one possible design, in another implementation of another aspect of the embodiments of this application, the comparison unit is specifically used for:

[0023] In the order described above, the target binary codes of multiple target character segments are compared with the candidate binary codes of multiple candidate character segments in the first document, wherein the first document is any one of the multiple candidate documents;

[0024] If the target binary code of any of the target character segments is different from the candidate binary code of the candidate character segments in the corresponding order, then a first comparison result is determined. The first comparison result is used to describe that the document to be processed is not a similar document to the first document.

[0025] or,

[0026] If the target binary code of each target character segment is the same as the candidate binary code of the candidate character segments in the corresponding order, then a second comparison result is determined. The second comparison result is used to describe that the document to be processed is a similar document to the first document.

[0027] In one possible design, in another implementation of another aspect of the embodiments of this application, the document processing apparatus further includes a storage unit. Specifically, the comparison unit is further used for:

[0028] If the target binary code of any target character segment is different from the candidate binary code of the candidate character segments in the corresponding order, then after determining the first comparison result, the target binary codes of multiple target character segments are compared with the candidate binary codes of multiple candidate character segments in the second document in the order stated above. The second document is any document other than the first document among the multiple candidate documents.

[0029] If the target binary code of any of the target character segments is different from the candidate binary code of the candidate character segments in the corresponding order in the second document, then a third comparison result is determined. The third comparison result is used to describe that the document to be processed and the second document are not similar documents.

[0030] The storage unit is specifically used to store the document to be processed in the candidate document library.

[0031] In one possible design, in another implementation of another aspect of the embodiments of this application, the document processing unit further includes a configuration unit, a processing unit, and a filtering unit. Specifically, the configuration unit is configured to, after determining the second comparison result, set a first identifier for the document to be processed and the first document if the target binary code of each target character segment is identical to the candidate binary code of the candidate character segments in the corresponding order. The first identifier is used to indicate that the document to be processed and the first document are similar documents.

[0032] The processing unit is specifically configured to, based on the first identifier, not save the document to be processed in the candidate document library;

[0033] or,

[0034] The filtering unit is specifically used to filter the document to be processed based on the first identifier.

[0035] In one possible design, in another implementation of another aspect of the embodiments of this application, the determining unit is specifically used for:

[0036] Extract the tags of each candidate domain lexicon from the multiple candidate domain lexicons. The tags of each candidate domain lexicon are used to describe the semantic domain to which the corresponding candidate domain lexicon belongs.

[0037] The document tags of the document to be processed are matched with the tags of each of the candidate domain thesauruses, so that the candidate domain thesauruses with matching tags are used as the target domain thesauruses matched by the document to be processed.

[0038] In one possible design, in another implementation of another aspect of the embodiments of this application, the target domain lexicon includes multiple candidate words and a word weight for each candidate word; the determining unit is specifically used for:

[0039] The document to be processed is segmented into multiple words, and each word is matched with multiple candidate words in the target domain thesaurus to determine the target candidate word that matches each word segment.

[0040] The word weight of each target candidate word is used as the word weight of the corresponding text segmentation.

[0041] In one possible design, in another implementation of another aspect of the embodiments of this application, the document processing unit further includes a generation unit and an assignment unit. Specifically, the acquisition unit is further configured to: acquire the candidate document library, which includes multiple candidate documents, before determining the target domain thesaurus matched by the document to be processed from multiple candidate domain thesauruses based on the document tags of the document to be processed;

[0042] The determining unit is specifically used to extract the document tags of each candidate document, and the document tags of each candidate document are used to describe the semantic domain to which the corresponding candidate document belongs;

[0043] The determining unit is specifically used to determine the target candidate word in each of the candidate documents, wherein the target candidate word is a word whose frequency of occurrence in the corresponding candidate document is greater than a preset threshold;

[0044] The generation unit is specifically used to classify the target candidate words of multiple candidate documents based on the document tags of each candidate document, and generate multiple candidate domain lexicons, each of which is associated with a semantic domain;

[0045] The assignment unit is specifically used to assign weights to target candidate words in each candidate domain thesaurus based on the document influence of each target candidate word in the corresponding candidate document, so as to obtain the word weights of each target candidate word in each candidate domain thesaurus.

[0046] In one possible design, in another implementation of another aspect of the embodiments of this application, the determining unit is specifically used for:

[0047] Text features are extracted from multiple text segments of the document to be processed to obtain the text feature vector of each text segment in the document to be processed.

[0048] Based on the word weight of each text segment, the text feature vector of the corresponding text segment is weighted to obtain the weighted feature vector of each text segment;

[0049] The weighted feature vectors of all the text segmentation words in the document to be processed are merged to obtain the fusion feature of the document to be processed.

[0050] The fused features are subjected to feature dimensionality reduction processing to obtain the target hash signature of the document to be processed.

[0051] In one possible design, in another implementation of another aspect of the embodiments of this application, the determining unit is specifically used to: calculate the hash value of each of the text segments of the document to be processed, and obtain the text feature vector corresponding to the text segment.

[0052] In one possible design, in another implementation of another aspect of the embodiments of this application, the determining unit is specifically used for:

[0053] Extract the document title and text content of the document to be processed;

[0054] Based on a preset document classification model, the document title and text content of the document to be processed are classified to obtain the document tags of the document to be processed.

[0055] A third aspect of this application provides a document processing apparatus, including: a memory, an input / output interface, and a processor. The memory stores program instructions. The processor executes the program instructions in the memory to perform the document processing method corresponding to the embodiments of the first aspect described above.

[0056] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method corresponding to the embodiments of the first aspect described above.

[0057] The fifth aspect of this application provides a computer program product containing instructions that, when run on a computer or processor, causes the computer or processor to execute the method described above for performing the implementation method of the first aspect.

[0058] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0059] In this embodiment, the document to be processed is first acquired, and its document tags are determined. These document tags describe the semantic domain to which the document belongs. Additionally, the document to be processed is segmented into multiple text segments. Then, based on the document tags, a target domain thesaurus matching the document is determined from multiple candidate domain thesauruses, and the word weight of each text segment is determined from the target domain thesaurus. The word weights describe the importance of the corresponding text segment in the semantic domain. After determining the word weights of each text segment, a target hash signature is determined based on the multiple text segments and their word weights. The target hash signature of the document to be processed is then compared with the hash signatures of multiple candidate documents in the candidate document library to determine the comparison result. This comparison result reflects whether the document to be processed and the candidate documents are similar documents. Through this method, this application pre-sets multiple candidate domain thesauruses based on a large number of documents, and different candidate domain thesauruses correspond to different semantic domains. In this way, after classifying the documents to be processed according to document tags, this application can determine the word weights of each text segment in the document based on the domain lexicon of the semantic domain. Then, it performs subsequent word segmentation weighting, hash signature calculation, and comparison according to the word weights of its respective domain. Thus, even if documents have highly overlapping word segments, the calculated hash signatures will be different due to the different word weights of their semantic domains. This effectively avoids misclassifying a large number of documents under the same category as similar, reduces the probability of misclassification, significantly improves the efficiency and accuracy of comparison, and reduces the loss of document information. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 This application provides a schematic diagram of an application scenario.

[0062] Figure 2 A schematic diagram of an implementation environment in which the document processing method provided in this application can be applied is shown;

[0063] Figure 3 A flowchart of a document processing method provided in an embodiment of this application is shown;

[0064] Figure 4This document illustrates a framework flowchart for determining document tags according to an embodiment of the present application.

[0065] Figure 5 This paper illustrates a framework flowchart for determining a target hash signature according to an embodiment of this application.

[0066] Figure 6 An optional schematic diagram illustrating the determination of the target hash signature provided in an embodiment of this application is shown;

[0067] Figure 7 This illustration shows a framework diagram for configuring word weights in different semantic domains, provided in an embodiment of this application.

[0068] Figure 8 This paper illustrates a flowchart of a process for determining the comparison result provided in an embodiment of this application.

[0069] Figure 9 This illustration shows a framework diagram for determining the comparison result based on binary code, provided in an embodiment of this application.

[0070] Figure 10 An optional schematic diagram of binary code comparison provided in an embodiment of this application is shown;

[0071] Figure 11 A schematic diagram of the functional module structure of the document processing apparatus provided in an embodiment of this application is shown.

[0072] Figure 12 A schematic diagram of the hardware structure of the document processing device provided in the embodiments of this application is shown. Detailed Implementation

[0073] This application provides a document processing method, apparatus, device, storage medium, and program product, which effectively avoids a large number of documents under the same category being misjudged as similar, reduces the probability of misjudgment, greatly improves the efficiency and accuracy of comparison, and reduces the loss of document information.

[0074] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0075] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0076] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that implementations of the application described herein can be implemented, for example, in sequences other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0077] In today's rapidly developing information technology landscape, document processing has become an essential part of daily life and work, and a key force driving innovation across various industries. While different documents may differ in content, structure, and format, they may contain repetitions or overlaps in high-frequency words and phrases. A large number of duplicate or similar documents can lead to poor information retrieval efficiency and higher storage requirements. Therefore, appropriate methods for document deduplication are necessary to improve information retrieval efficiency.

[0078] For example, Figure 1 A schematic diagram illustrating an application scenario provided in this application is shown. For example... Figure 1 As shown, when document A is obtained, to determine whether document A can be included in the document library, it is necessary to check whether other documents similar to document A exist in the document library. For example, if the document library contains documents B, C, and D, document A needs to be compared with documents B, C, and D in turn for similarity. If the similarity between document A and document B is 80%, which is considered highly similar, then document A can be filtered out and not stored in the document library, thus achieving deduplication of document A. This ensures that there are not a large number of duplicate documents in the document library.

[0079] However, current deduplication schemes typically use Locality Sensitive Hashing (SIMH) to calculate the Hamming distance between the hash signatures of multiple documents, and then use this Hamming distance to determine whether the compared documents are similar. However, in cases where high-frequency words in some documents may highly overlap, if current deduplication techniques use SIMH to calculate the Hamming distance for these high-frequency words, the calculated distances will be extremely close. For example, certain technical terms in a legal document A may highly overlap with certain terms in an academic paper B. Even if legal document A and academic paper B belong to different semantic domains, current deduplication techniques will still calculate that the Hamming distance between legal document A and academic paper B is close, thus judging legal document A and academic paper B as similar documents. Therefore, according to current deduplication schemes, documents with close Hamming distances may be mistakenly identified as similar and subjected to deduplication. This not only increases the probability of misidentification but may also lead to problems such as loss of document information after incorrect deduplication.

[0080] Therefore, to solve the aforementioned technical problems, this application provides a document processing method. This document processing method can be applied to scenarios such as text deduplication. For example, it can be applied to information retrieval, data cleaning, copyright protection, and information security, etc., without specific limitations in this application. In the document processing method of this application, multiple candidate domain thesauruses are pre-set based on a large number of documents, and different candidate domain thesauruses correspond to different semantic domains. Thus, after classifying the documents to be processed according to document tags, this application can determine the word weights of each text segment in the document based on the domain thesaurus of its respective semantic domain, and then perform subsequent word segmentation weighting, hash signature calculation, and comparison processing according to the word weights of their respective domains. In this way, even if documents have highly overlapping word segments, the calculated hash signatures will not be the same due to the different word weights of their respective semantic domains. This effectively avoids a large number of documents under the same category being misjudged as similar, reduces the probability of misjudgment, significantly improves the efficiency and accuracy of comparison, and reduces the loss of document information.

[0081] For example, the document processing method provided in this application can be applied to one or more of the following application scenarios, as follows:

[0082] Scenario 1: Information Retrieval

[0083] Applications such as search engines and document management systems that need to process massive amounts of text data typically use document deduplication techniques to improve search efficiency and reduce redundant information. Applying the document processing method provided in this application to this information retrieval field allows for the comprehensive consideration of document tags—that is, the influence of the semantic domain on word segmentation—before calculating the hash signatures of different documents. Thus, based on the domain lexicon of the semantic domain, the word weights of each text segment in the document to be processed are determined, and subsequent word segmentation weighting, hash signature calculation, and comparison are performed according to the word weights of their respective domains. Even when documents have highly overlapping word segments, the different word weights of their semantic domains will result in different final hash signatures. This effectively avoids misclassifying a large number of documents in the same category as similar, reduces the probability of misclassification, significantly improves the efficiency and accuracy of comparison, helps accelerate the indexing of search engines, and reduces the interference of redundant documents on search results.

[0084] Scenario 2: Data Cleaning

[0085] In the process of data cleaning, situations often arise where there is a large amount of duplicate information in the data. Applying the document processing method provided in this application to the field of information retrieval allows for the comprehensive consideration of document tags of different documents before calculating the hash signatures of different documents, that is, the influence of the semantic domain of the document on word segmentation. Thus, based on the domain lexicon of the semantic domain, the word weights of each text segment in the document to be processed are determined, and subsequent word segmentation weighting, hash signature calculation, and comparison are performed according to the word weights of their respective domains. Even when documents have highly overlapping word segments, the different word weights of their semantic domains will result in different final hash signatures. This effectively avoids a large number of documents in the same category being misclassified as similar, reduces the probability of misclassification, significantly improves the efficiency and accuracy of comparison, and enhances data quality. This is of great significance for data analysis, machine learning model training, and other fields.

[0086] Scenario 3: Natural Language Processing

[0087] In natural language processing (NLP) scenarios, a large number of duplicate documents can lead to poor generalization ability of model training. Therefore, applying the document processing method provided in this application to this NLP field can comprehensively consider the document tags of different documents, that is, the influence of the semantic domain of the document on word segmentation. Thus, based on the domain lexicon of the semantic domain, the word weights of each text segment in the document to be processed are determined, and subsequent word segmentation weighting, hash signature calculation, and comparison are performed according to the word weights of their respective domains. In this way, even if there are documents with highly overlapping word segments, the hash signatures calculated will not be the same because the word weights of their respective semantic domains are different. This effectively avoids a large number of documents in the same category being misclassified as similar, reduces the probability of misclassification, significantly improves the efficiency and accuracy of comparison, helps to clean the corpus needed for model training, and improves the generalization ability of the model.

[0088] It should be noted that, in addition to the information retrieval, data cleaning, and natural language processing scenarios mentioned above, the document processing method of this application can also be applied to other document deduplication scenarios in practice. For example, it can also be applied to the fields of information security and copyright protection, etc., without specific limitations in this application.

[0089] For example, the document processing method provided in this application can be applied to Figure 2 The implementation environment shown is where the document processing device 100 performs operations. Figure 2 The illustrated implementation environment includes a document processing device 100 and, for example, a user terminal 101. The document processing device 100 and the user terminal 101 can be directly or indirectly connected via wired or wireless communication, etc., without specific limitations in this application. Optionally, the implementation environment may also include a database 102, etc. The document processing device 100 can also be connected to the database 102 via a network.

[0090] like Figure 2 As shown, in step S1, the user can send the document to be processed to the document processing device 100 through the user terminal 101. For example, after the user inputs the document to be processed, the user terminal 101 carries the document to be processed in the processing request, and then sends the processing request to the document processing device 100, thereby the document processing device 100 obtains the document to be processed carried in the processing request.

[0091] Furthermore, in step S2, the document processing device 100 determines the document tags of the document to be processed. It should be noted that the document tags of the document to be processed are used to describe the semantic domain to which the document belongs. In step S3, the document processing device 100 performs word segmentation on the document to be processed, obtaining multiple text segments of the document. In step S4, based on the document tags of the document to be processed, the document processing device 100 determines the target domain lexicon matching the document to be processed from multiple candidate domain lexicons, and determines the word weight of each text segment of the document to be processed from the target domain lexicon. The described word weight is used to describe the importance of the corresponding text segment in the semantic domain. For example, the document processing device 100 can obtain multiple candidate domain lexicons from storage media such as database 102; this application does not limit this.

[0092] In step S5, the document processing device 100 determines the target hash signature of the document to be processed based on multiple text segmentations of the document and the word weight of each text segmentation in the document. In step S6, the document processing device 100 compares the target hash signature of the document to be processed with the hash signatures of multiple candidate documents in the candidate document library to determine the comparison result. The comparison result is used to reflect whether the document to be processed and the candidate documents are similar documents.

[0093] Optionally, in Figure 2 The implementation environment shown may also include step S7. That is, in step S7, the document processing device 100 may also send the comparison result to the user terminal 101, and then the user terminal 101 may display the comparison result through a client, engine or other visual interface to prompt the user whether the document to be processed and the candidate document are similar documents.

[0094] The user terminal 101 mentioned in this application can be a processing device with data processing capabilities, such as a terminal device. The terminal device may include, but is not limited to, smartphones, desktop computers, laptops, tablets, smart speakers, in-vehicle devices, smartwatches, wearable smart devices, and smart voice interaction devices.

[0095] The document processing device 100 mentioned in this application can be a processing device with data processing capabilities, such as a terminal device or a server. The terminal device can include, but is not limited to, smartphones, desktop computers, laptops, tablets, smart speakers, in-vehicle devices, smartwatches, wearable smart devices, and smart voice interaction devices. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. This application does not specifically limit the scope of the application.

[0096] The database 102 mentioned in this application can be simply viewed as an electronic filing cabinet—a place to store electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, shared by multiple users, with minimal redundancy, and independent of application programs. A Database Management System (DBMS) is a computer software system designed to manage databases, generally possessing basic functions such as storage, retrieval, security, and backup. DBMSs can be classified according to the database model they support, such as relational or Extensible Markup Language (XML); or according to the type of computer they support, such as server clusters or mobile phones; or according to the query language used, such as Structured Query Language (SQL) or XQuery; or according to performance priorities, such as maximum scale or maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, simultaneously supporting multiple query languages. Database 102 in this application can be used to store pre-trained language models, etc. Optionally, it can also be used to store documents to be processed, multiple candidate domain lexicons, etc.

[0097] For example, the document processing method provided in this application, in addition to those described above... Figure 2 Beyond the application scenarios mentioned, in practical applications, it can also be applied to fields such as artificial intelligence (AI), blockchain, and cloud technology, which are not limited in this application.

[0098] The following describes a document processing method provided by an embodiment of this application with reference to the accompanying drawings. Figure 3 A flowchart illustrating a document processing method provided in an embodiment of this application is shown, which is described only as an example using a document processing device as the execution subject. Figure 3 As shown, the document processing method may include the following steps:

[0099] 301. Obtain the document to be processed and determine the document tags of the document to be processed. The document tags of the document to be processed are used to describe the semantic domain to which the document to be processed belongs.

[0100] In one or more embodiments, the document to be processed can be understood as a document that has not yet undergone deduplication. For example, the document to be processed includes, but is not limited to, documents that have not undergone deduplication and are not saved in the document library, or other documents that have not undergone deduplication but are saved in the document library, etc., which are not limited in this application. For example, the document to be processed could be "I like watching the Brazil World Cup", or "I like to travel to the Arctic in winter", etc., which are not limited in this application.

[0101] As an illustrative example, if an object wishes to determine whether a document to be processed is similar to other documents, it can input the document to be processed using a user terminal and then send the document to be processed to a document processing device via a communication network. In this way, the document processing device obtains the document to be processed. Alternatively, if the document to be processed is pre-stored in a database, the document processing device can also retrieve the document from the database. This application does not specifically limit the method of retrieving the document to be processed.

[0102] After acquiring the document to be processed, this application also needs to perform tag recognition on the document to determine its document tags. The document tags of the document to be processed can be used to identify the semantic domain to which the document belongs. For example, for a Chinese language test paper in a document library scenario, its document classification can be defined as follows: First-level classification: Educational Examinations, Second-level classification: Basic Education, Third-level classification: Subjects, Fourth-level classification: Chinese Language.

[0103] Optionally, in the foregoing Figure 3 Based on one or more of the described embodiments, in another embodiment provided in this application, regarding how to determine the document tags of the document to be processed in step 301, it can be referred to Figure 4 Use the schematic diagram shown to understand the framework.

[0104] like Figure 4As shown, in determining the document tags of the document to be processed, the document title and text content of the document to be processed can be extracted first. For example, since document titles usually have obvious formatting features, such as bold font, color, and position, a title extraction tool can be used to extract the corresponding document title from the document to be processed. Alternatively, a text extraction tool such as a text editor can be used to extract the text content from the document to be processed; this application does not impose specific limitations. It should be noted that the described text content can be understood as the main body of the document to be processed, containing text information, data, etc. The described document title can be used to describe the subject of the document to be processed.

[0105] In this way, after extracting the document title and text content of the document to be processed, a pre-defined document classification model can be used to identify document tags. As an illustrative example, the document title and text content of the document to be processed are input into the pre-defined document classification model. The model then performs category recognition processing on the document title and text content, thereby identifying the document tags of the document to be processed. For example, the pre-defined document classification model calculates the probability of each tag based on the input document title and text content, and then selects the tag with the highest probability as the document tag of the document to be processed. By using a pre-defined document classification model to determine the document tags of the document to be processed in this way, the time and effort required for manual classification are significantly reduced, the overall efficiency of document processing is improved, and it is easy to quickly locate the semantic domain to which the document belongs.

[0106] It should be noted that the preset document classification model involved in this application may include, but is not limited to, neural network models such as convolutional neural network (CNN) and recurrent neural network (RNN), but no specific limitation is made in this application.

[0107] 302. Perform word segmentation on the document to be processed to obtain multiple text segments of the document.

[0108] In one or more embodiments, after obtaining the document to be processed, it is also necessary to perform word segmentation on the document to obtain multiple text segments of the document. The word segmentation described in this application is the process of recombinizing the continuous sequence of characters corresponding to the text into a word sequence according to certain rules.

[0109] As an illustrative example, each string in the document to be processed can be matched against strings in a pre-defined dictionary. If a string matches a string in the pre-defined dictionary, the string in the document is segmented as a word, thus obtaining multiple text segments of the document. Alternatively, the characters in the document can be converted into corresponding character vectors, which are then input into a neural network-based machine learning model to obtain the probability that the corresponding character belongs to a pre-defined word position labeling state. Based on the probability, the word position labeling state of each character in the document is determined, and the document is then segmented according to the word position labeling state of each character, thus obtaining multiple text segments.

[0110] It should be noted that the preset position annotation states in each word mentioned can be understood as the preset position annotations corresponding to the current character's position within its word. For example, position annotation B indicates that the character is at the beginning of the word, position annotation M indicates that the character is in the middle of the word, position annotation E indicates that the character is at the end of the word, and position annotation S indicates that the character is a word on its own, etc. This application does not impose specific limitations. Furthermore, the neural network-based machine learning model can be a memory network model, such as a long short-term memory (LSTM) network, a bidirectional long short-term memory (Bi-LSTM) network, or an RNN neural network, etc. This application does not impose specific limitations.

[0111] For example, if the document to be processed is “similarity calculation of massive data based on file classification”, by segmenting the document into words, multiple corresponding text words can be obtained, such as “based on”, “file”, “classification”, “of”, “massive”, “data”, “similarity”, “calculation”, etc.

[0112] 303. Based on the document tags of the document to be processed, determine the target domain thesaurus that the document to be processed matches from multiple candidate domain thesauruses, and determine the word weight of each text segment of the document to be processed from the target domain thesaurus. The word weight is used to describe the importance of the corresponding text segment in the semantic domain.

[0113] In one or more embodiments, each candidate domain lexicon is associated with a semantic domain. For example, candidate domain lexicon A is associated with semantic domain A, and each candidate word in candidate domain lexicon A belongs to semantic domain A; candidate domain lexicon B is associated with semantic domain B, and each candidate word in candidate domain lexicon B belongs to semantic domain B. The word weights of candidate words in different semantic domains are different. For example, the word weights of each candidate word in candidate domain lexicon A are different from the word weights of each candidate word in candidate domain lexicon B. For instance, if a word 'a' belongs to both semantic domain B and semantic domain A, then the word weight of word 'a' in semantic domain A is different from the word weight of word 'a' in semantic domain B.

[0114] Based on this, after determining the document tags of the document to be processed, the target domain thesaurus matching the document is determined from multiple candidate domain thesauruses according to these tags. As an illustrative description, in determining the target domain thesaurus, the tags of each candidate domain thesaurus can be extracted first. Each candidate domain thesaurus's tag describes the semantic domain to which it belongs. Further, the document tags of the document to be processed are matched with the tags of each candidate domain thesaurus. Thus, the candidate domain thesaurus with matching tags is taken as the target domain thesaurus matching the document to be processed. Through this method, by matching document tags with the tags of each candidate domain thesaurus, the domain thesaurus most relevant to the content of the document to be processed can be accurately found. This method avoids the subjectivity and uncertainty of manual judgment, improves the accuracy of matching, and lays the foundation for subsequent accurate matching and locating the word weights of each word segment in the document to be processed.

[0115] Therefore, after determining the target domain thesaurus that matches the document to be processed from multiple candidate domain thesauruses, this application also needs to determine the word weight of each text segment of the document to be processed from the target domain thesaurus. For example, since the target domain thesaurus includes multiple candidate words and the word weight of each candidate word, after determining the target domain thesaurus, the multiple text segments of the document to be processed can be matched with the multiple candidate words in the target domain thesaurus respectively, thereby determining the target candidate words that match each text segment. Then, the word weight of each target candidate word is used as the word weight of the corresponding text segment. Through the above method, matching text segments with words in the target domain thesaurus helps to quickly and accurately determine the word weight of the text segment in its domain, further improving the accuracy of subsequent hash signature calculations.

[0116] 304. Based on multiple text segmentations of the document to be processed and the word weights of each text segmentation in the document to be processed, determine the target hash signature of the document to be processed.

[0117] In one or more embodiments, after obtaining the word weight of each text segment in the document to be processed using a candidate domain lexicon, the word weight of each text segment can be weighted and applied to the corresponding text segment to determine the target hash signature of the document to be processed. The authenticity and integrity of the document to be processed can be verified through the target hash signature.

[0118] As an illustrative description, the process of determining the target hash signature of the document to be processed in step 304 can be referred to as follows: Figure 5 Use the illustrated framework flowchart to understand the process.

[0119] like Figure 5 As shown, in the process of determining the target hash signature based on text segmentation and word weights, text features can be extracted from multiple text segments of the document to be processed to obtain the text feature vector of each text segment in the document. For example, the text feature vector of the corresponding text segment can be obtained by calculating the hash value of each text segment of the document.

[0120] Thus, after calculating the text feature vector of each text segment in the document to be processed, the text feature vectors of the corresponding text segments are weighted based on the word weight of each text segment to obtain a weighted feature vector for each text segment. Then, the weighted feature vectors of all text segments in the document to be processed are merged to obtain the fused feature of the document. Finally, feature dimensionality reduction processing is performed on the fused feature to obtain the target hash signature of the document to be processed.

[0121] The above method first applies word weights to the text feature vectors, making the text representation more accurate. Then, the weighted feature vectors are merged and subjected to dimensionality reduction to obtain the document's hash signature. This reduces data dimensionality, not only decreasing storage requirements but also improving computational efficiency, thus enhancing the efficiency and accuracy of subsequent text deduplication calculations.

[0122] For example, Figure 6 An optional schematic diagram of determining the target hash signature provided in an embodiment of this application is shown. For example... Figure 6As shown, taking the document to be processed as "Massive Data Similarity Calculation Based on File Classification" as an example, by segmenting the document "Massive Data Similarity Calculation Based on File Classification", we can obtain multiple corresponding text segments, such as "based on", "file", "classification", "of", "massive", "data", "similarity", and "calculation". By performing hash calculations on the segments such as "based on", "file", "classification", "of", "massive", "data", "similarity", and "calculation", we obtain the text feature vectors of the segments "based on" as "100101", "file" as "101011", "classification" as "101010", "of" as "001010", "...", "data" as "101100", "similarity" as "011001", and "calculation" as "110110".

[0123] If the document to be processed belongs to semantic domain A, and the word weights of "based on", "file", "category", "of", ..., "data", "similarity", and "calculation" are 3, 3, 4, 1, ..., 6, 2, and 5 respectively, then after weighting, the weighted feature vectors are: "based on", "file", "category", "of", ..., "data", "similarity", and "calculation".

[0124] Thus, after merging, the fused feature is "36,-24,7,-9,10,12". Subsequently, the dimensionality of the fused feature is reduced to obtain the target hash signature of the document to be processed as "1,-1,1,-1,1,1".

[0125] Optionally, after calculating the target hash signature of the document to be processed, the target hash signature can also be stored.

[0126] 305. Compare the target hash signature of the document to be processed with the hash signatures of multiple candidate documents in the candidate document library to determine the comparison result. The comparison result is used to reflect whether the document to be processed and the candidate documents are similar documents.

[0127] In one or more embodiments, after calculating the target hash signature of the document to be processed, the target hash signature of the document to be processed can be compared with the hash signature of each candidate document in the candidate document library to determine the comparison result. In this way, the comparison result can clearly determine whether the document to be processed and the candidate documents are similar documents. For example, if they are similar documents, filtering processing can be performed to filter the document to be processed. Conversely, if they are not similar documents, the document to be processed can be saved to the candidate document library.

[0128] By employing the above method, this application pre-sets multiple candidate domain thesauruses based on a large number of documents, with each candidate domain thesaurus corresponding to a different semantic domain. Thus, after classifying the documents to be processed according to document tags, this application can determine the word weights of each text segment in the document based on the domain thesaurus of its respective semantic domain. Subsequent processing, such as word segmentation weighting, hash signature calculation, and comparison, is then performed according to the word weights of their respective domains. Therefore, even when documents have highly overlapping word segments, the calculated hash signatures will differ due to the different word weights within their semantic domains. This effectively avoids misclassifying a large number of documents under the same category as similar, reducing the probability of misclassification, significantly decreasing the volume of document comparisons, greatly improving the efficiency and accuracy of comparisons, and minimizing the loss of document information.

[0129] Optionally, in the foregoing Figure 3 Based on one or more corresponding embodiments, before determining the word weights of each text segmentation in step 303, this application can also pre-set different word weights for words in different semantic domains, as can be referred to... Figure 7 Use the schematic diagram shown to understand the framework.

[0130] like Figure 7 As shown, before performing the aforementioned step 303, the document processing device may first acquire a candidate document library. This candidate document library includes multiple candidate documents. After acquiring the candidate document library, the document tags for each candidate document in the library can be extracted.

[0131] It should be noted that the document tags for each candidate document describe the semantic domain to which the corresponding candidate document belongs. For example, taking any candidate document (such as the first candidate document) as an example, to extract the document tags for the first candidate document, we can first extract the document title and text content of the first candidate document. After extracting the document title and text content of the first candidate document, we can use a pre-defined document classification model to identify the document tags. For example, we can input the document title and text content of the first candidate document into the pre-defined document classification model, and use the pre-defined document classification model to perform category recognition processing on the document title and text content of the first candidate document, thereby identifying the document tags of the first candidate document. For details, please refer to the aforementioned... Figure 4 The framework flowchart shown is for reference only and will not be elaborated upon here.

[0132] Furthermore, after obtaining each candidate document in the candidate document library, target candidate words can be determined for each candidate document. For example, taking any candidate document (such as the first candidate document), the frequency of each candidate word in the first candidate document is counted, and then the frequency of occurrence of the corresponding candidate word is calculated based on the frequency of occurrence of each candidate word. In this way, words with a frequency greater than a certain preset threshold are selected as target candidate words in the first candidate document.

[0133] In this way, based on the document tags of each candidate document, the target candidate words of multiple candidate documents are classified to generate multiple candidate domain lexicons. For example, candidate documents with the same document tags are grouped into one category, and then a candidate domain lexicon corresponding to that document tag is constructed based on the target candidate words in one or more candidate documents with the same document tags. It should be noted that each candidate domain lexicon is associated with a semantic domain. Finally, the influence of each target candidate word in its corresponding candidate document is evaluated to obtain the document influence of each target candidate word in its corresponding candidate document. Thus, based on the document influence of each target candidate word in its corresponding candidate document, the target candidate words in each candidate domain lexicon are weighted to obtain the word weight of each target candidate word in each candidate domain lexicon. Optionally, special words can be freely customized into the candidate domain lexicon according to deduplication requirements, etc., and weights can be configured for them to provide a basis for dimensionality reduction processing in the subsequent hash signature calculation process.

[0134] It should be noted that the word weight of each target candidate word can be used to describe the importance of the corresponding target candidate word in the corresponding semantic domain.

[0135] By using the above method, different word weights are configured for each candidate word in different semantic domains, so that even documents under the same category can calculate different target hash signatures because they belong to different semantic domains. This provides correct data support for the weighted calculation in the subsequent document deduplication process, which is conducive to improving the deduplication effect efficiently and accurately.

[0136] In some other alternative embodiments, in the foregoing Figure 3 Based on one or more embodiments, the process for determining the comparison result in step 305 can be referred to the following: Figure 8 Use the illustrated framework flowchart to understand the process. Figure 8 As shown, it includes at least the following steps:

[0137] S3051. Split the target hash signature of the document to be processed into multiple target character segments arranged in order, each target character segment including the target binary code.

[0138] For example, suppose the target hash signature of the document to be processed is a 16-bit string, such as "0000111122223333". This target hash signature can then be split into four target character segments arranged in sequence: target character segment 1, target character segment 2, target character segment 3, and target character segment 4. Each target character segment contains a 4-bit target binary code, such as: target character segment 1 contains "0000", target character segment 2 contains "1111", target character segment 3 contains "2222", and target character segment 4 contains "3333".

[0139] S3052. The hash signature of each candidate document is split into strings to obtain multiple candidate character segments arranged in order for each candidate document. Each candidate character segment includes candidate binary code.

[0140] For example, if multiple candidate documents include candidate document 1, candidate document 2, and candidate document 3, and assuming that the hash signature of candidate document 1 is a 16-bit string, such as "0000333344445555", after splitting, we can obtain the candidate binary codes of the candidate character segments in candidate document 1. For example, the candidate binary code included in candidate character segment 11 is "0000", the candidate binary code included in candidate character segment 12 is "3333", the candidate binary code included in candidate character segment 13 is "4444", and the candidate binary code included in candidate character segment 14 is "5555".

[0141] Similarly, assuming the hash signature of candidate document 2 is a 16-bit string, such as "1111222244445555", the hash signature of candidate document 2 can be split into four candidate character segments arranged in sequence, such as candidate character segment 21, candidate character segment 22, candidate character segment 23, and candidate character segment 24. Thus, each candidate character segment includes a 4-bit candidate binary code, for example: candidate character segment 21 contains the candidate binary code "1111", candidate character segment 22 contains the candidate binary code "2222", candidate character segment 23 contains the candidate binary code "4444", and candidate character segment 24 contains the candidate binary code "5555".

[0142] Similarly, assuming the hash signature of candidate document 3 is a 16-bit string, such as "0000111122223333", the candidate binary codes of the candidate character segments in candidate document 3 can be obtained in the same way as candidate document 1. For example, the candidate binary code included in candidate character segment 31 is "0000", the candidate binary code included in candidate character segment 32 is "1111", the candidate binary code included in candidate character segment 33 is "2222", and the candidate binary code included in candidate character segment 34 is "3333".

[0143] It should be noted that the target hash signature of the document to be processed mentioned above is a 16-bit string, such as "0000111122223333", and the hash signature of candidate document 1 is a 16-bit string, such as "0000333344445555". In practical applications, it can also be other numbers of strings, such as 64 bits, or it can be split into four 16-bit binary codes, etc. This application does not make specific limitations.

[0144] S3053. Based on the target binary codes of multiple target character segments in the document to be processed, and the candidate binary codes of multiple candidate character segments in each candidate document, determine the comparison result.

[0145] In one or more embodiments, after obtaining the target binary codes of multiple target character segments in the document to be processed, and the candidate binary codes of multiple candidate character segments in each candidate document, the target binary codes of multiple target character segments can be compared sequentially with the candidate binary codes of multiple candidate character segments in each candidate document to determine the comparison results.

[0146] Unlike related solutions that directly calculate the Hamming distance between the target hash signature of the document to be processed and the hash signature of the candidate document to perform document deduplication, this application performs string splitting on both the target hash signature of the document to be processed and the hash signature of the candidate document, and then calculates the similarity of documents by comparing binary codes in the same order. This not only reduces the massive data matching process and improves the efficiency of data comparison, but also ensures accuracy down to every byte in the document, preventing document misjudgment and improving the accuracy of document comparison.

[0147] In some alternative embodiments, regarding the foregoing Figure 8 In step S3053, how to determine the comparison result based on the target binary code and the candidate binary code, can be referred to... Figure 9 Use the flowchart shown to understand the framework.

[0148] like Figure 9 As shown, after obtaining the target binary codes of multiple target character segments in the document to be processed, and the candidate binary codes of multiple candidate character segments in each candidate document, the target binary codes of the multiple target character segments are compared sequentially with the candidate binary codes of the multiple candidate character segments in the first document. It should be noted that the first document is any one of the multiple candidate documents.

[0149] During the comparison process, two scenarios typically occur. Scenario 1: If the target binary code of any target character segment is different from the candidate binary code of the corresponding candidate character segments, then the first comparison result is determined. This first comparison result indicates that the document to be processed and the first document are not similar documents.

[0150] Conversely, in case 2, if the target binary code of each target character segment is the same as the candidate binary code of the corresponding candidate character segments, then the second comparison result is determined. That is, through this second comparison result, it can be described that the document to be processed and the first document are similar documents.

[0151] By comparing the target binary code with the candidate binary code sequentially as described above, if any binary code is different, the documents are determined to be dissimilar; otherwise, if they are identical, they are determined to be similar. This reduces the need for full data matching, enabling fast and efficient document comparison and improving both efficiency and accuracy.

[0152] Optionally, in the foregoing Figure 9Based on the illustrated embodiment, after determining the first comparison result for case 1, this application can continue to compare the target binary codes of multiple target character segments sequentially with the candidate binary codes of multiple candidate character segments in the second document, where the second document is any document other than the first document among the multiple candidate documents. If the target binary code of any target character segment is different from the candidate binary codes of the corresponding candidate character segments in the second document, a third comparison result is determined. The third comparison result is used to describe that the document to be processed and the second document are not similar documents. At this time, the document to be processed can be saved in the candidate document library. That is to say, after the above comparison process, if it is determined that each target binary code of the document to be processed is not completely identical to each candidate binary code of all candidate documents, it can be determined that the document to be processed is not a similar document to any of the candidate documents. At this time, the document to be processed can be saved in the candidate document library. Using the above method, even if the document to be processed is not similar to the first document, it is necessary to further compare the document to be processed with other documents by comparing their binary codes. This comparison method, because it does not require complex parsing and understanding of the document content, helps to quickly filter and locate similar documents from a massive amount of data, improving comparison efficiency. Furthermore, if all documents are determined to be dissimilar through binary code comparison, then the document is added to the database, which helps with subsequent storage and retrieval, ensuring the cleanliness and orderliness of the document database.

[0153] Optionally, in the foregoing Figure 9 Based on the illustrated embodiment, after determining the second comparison result in case 2, this application can further set a first identifier for the document to be processed and the first document. The first identifier is used to indicate that the document to be processed and the first document are similar documents. In this case, based on the first identifier, the document to be processed can be not saved in the candidate document library; or, based on the first identifier, the document to be processed can be filtered, avoiding the existence of documents with duplicate or similar content in the candidate document library, improving storage requirements, and increasing the efficiency of subsequent document retrieval, etc.

[0154] For example, Figure 10 An optional schematic diagram of binary code comparison provided in an embodiment of this application is shown. For example... Figure 10 As shown above, Figure 8 Taking the binary code of the document to be processed in step S3051 and the candidate document 1 in step S3052 as examples, the target binary code included in target character segment 1 of the document to be processed is "0000", the target binary code included in target character segment 2 is "1111", the target binary code included in target character segment 3 is "2222", and the target binary code included in target character segment 4 is "3333".

[0155] Taking the first document as the aforementioned candidate document 1 as an example, its candidate character segment 11 includes the candidate binary code "0000", candidate character segment 12 includes the candidate binary code "3333", candidate character segment 13 includes the candidate binary code "4444", and candidate character segment 14 includes the candidate binary code "5555".

[0156] At this point, following the sequence, the target binary code "0000" of target character segment 1 is first matched with the candidate binary code "0000" of candidate character segment 11. After comparison, "0000" is the same as "0000". Then, the target binary code "1111" of target character segment 2 is further compared with the candidate binary code "3333" of candidate character segment 12. After comparison, "1111" is different from "3333". At this point, it can be determined that the document to be processed is not similar to candidate document 1. Therefore, there is no need to match target character segment 3 with candidate character segment 13, or target character segment 4 with candidate character segment 14, reducing the cumbersome data matching process and improving comparison efficiency and effectiveness. In other words, as long as any one of these four target binary codes is different from any one of the four candidate binary codes in the corresponding sequence, it can be determined that the document to be processed is not similar to candidate document 1.

[0157] Furthermore, based on the determination that the document to be processed and candidate document 1 are not similar documents, it is also necessary to compare the binary codes of the document to be processed and candidate document 2. Specifically, since candidate document 1 and candidate document 2 have already been processed to determine document similarity by checking whether their binary codes are the same, that is, after comparison, it can be seen that the candidate binary code "0000" of candidate character segment 11 of candidate document 1 is different from the candidate binary code "1111" of candidate character segment 21 of candidate document 2. However, it has already been found that the target binary code "0000" of target character segment 1 of the document to be processed is the same as the candidate binary code "0000" of candidate character segment 11 of candidate document 1. Thus, it can be directly determined that the target binary code "0000" of target character segment 1 of the document to be processed is different from the candidate binary code "1111" of candidate character segment 21 of candidate document 2, that is, it is determined that the document to be processed and candidate document 2 are not similar documents. By using the above method, it is not necessary to perform a complete comparison process between the document to be processed and the candidate document 2, which saves the complicated comparison process, saves comparison time, and improves comparison efficiency.

[0158] Furthermore, it is necessary to determine whether the document to be processed and candidate document 3 are similar documents. That is, in sequence, if the target binary code "0000", the target binary code "1111", the target binary code "2222", and the target binary code "3333" of target character segment 1 are all the same as the candidate binary codes "0000", "1111", "2222", and "3333" of candidate document 3, respectively, then it can be determined that the document to be processed and candidate document 3 are similar documents.

[0159] Optionally, after calculating the target hash signature of the document to be processed, if the document to be processed is not similar to any of the candidate documents after comparison, the target hash signature can be stored. For example, the target hash signature of the document to be processed can be stored in a candidate document library.

[0160] By employing the aforementioned method, this application splits the target hash signature of the document to be processed and the hash signature of the candidate document into strings, and then calculates the similarity of documents by comparing binary codes in the same order. This not only reduces the massive data matching process and improves data comparison efficiency, but also avoids misclassification of documents and improves the accuracy of document comparison. Furthermore, this application's scheme can perform noise reduction processing on documents within the same semantic domain, avoiding misclassification caused by high overlap of high-frequency words, thus reducing the probability of similarity errors in document entry. Moreover, this application's scheme is flexible, allowing the calculation results of hash signatures to differ depending on the document classification.

[0161] The foregoing primarily describes the solutions provided by the embodiments of this application from a methodological perspective. It is understood that to achieve the above functions, corresponding hardware structures and / or software modules are included to execute each function. Those skilled in the art should readily recognize that, based on the modules and algorithm steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0162] This application embodiment can divide the device into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0163] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0164] The document processing apparatus in the embodiments of this application will now be described in detail. Figure 11 This is a schematic diagram of the functional module structure of the document processing device provided in the embodiments of this application. For example... Figure 11 As shown, the document processing apparatus may include:

[0165] Acquisition unit 1101 is used to acquire the document to be processed;

[0166] The determining unit 1102 is used to determine the document tags of the document to be processed, wherein the document tags of the document to be processed are used to describe the semantic domain to which the document to be processed belongs;

[0167] The word segmentation unit 1103 is used to segment the document to be processed into multiple text words of the document to be processed.

[0168] The determining unit 1102 is configured to determine the target domain thesaurus that matches the document to be processed from multiple candidate domain thesauruses based on the document tags of the document to be processed, and to determine the word weight of each text segment of the document to be processed from the target domain thesaurus, wherein the word weight is used to describe the importance of the corresponding text segment in the semantic domain;

[0169] The determining unit 1102 is used to determine the target hash signature of the document to be processed based on multiple text segments of the document to be processed and the word weight of each text segment in the document to be processed;

[0170] The comparison unit 1104 is used to compare the target hash signature of the document to be processed with the hash signatures of multiple candidate documents in the candidate document library to determine the comparison result. The comparison result is used to reflect whether the document to be processed and the candidate documents are similar documents.

[0171] Optionally, in the above Figure 11 Based on the corresponding embodiments, in another embodiment of the document processing apparatus provided in this application, the comparison unit 1104 is specifically used for:

[0172] The target hash signature of the document to be processed is split into multiple target character segments arranged in order, and each target character segment includes target binary code;

[0173] The hash signature of each candidate document is split into strings to obtain multiple candidate character segments of each candidate document arranged in the order stated above, and each candidate character segment includes candidate binary code;

[0174] The comparison result is determined based on the target binary codes of multiple target character segments in the document to be processed and the candidate binary codes of multiple candidate character segments in each candidate document.

[0175] Optionally, in the above Figure 11 Based on one or more corresponding embodiments, in another embodiment of the document processing apparatus provided in this application, the comparison unit 1104 is specifically used for:

[0176] In the order described above, the target binary codes of multiple target character segments are compared with the candidate binary codes of multiple candidate character segments in the first document, wherein the first document is any one of the multiple candidate documents;

[0177] If the target binary code of any of the target character segments is different from the candidate binary code of the candidate character segments in the corresponding order, then a first comparison result is determined. The first comparison result is used to describe that the document to be processed is not a similar document to the first document.

[0178] or,

[0179] If the target binary code of each target character segment is the same as the candidate binary code of the candidate character segments in the corresponding order, then a second comparison result is determined. The second comparison result is used to describe that the document to be processed is a similar document to the first document.

[0180] Optionally, in the above Figure 11Based on one or more corresponding embodiments, in another embodiment of the document processing apparatus provided in this application, the document processing apparatus further includes a storage unit 1105. The comparison unit 1104 is specifically further used for:

[0181] If the target binary code of any target character segment is different from the candidate binary code of the candidate character segments in the corresponding order, then after determining the first comparison result, the target binary codes of multiple target character segments are compared with the candidate binary codes of multiple candidate character segments in the second document in the order stated above. The second document is any document other than the first document among the multiple candidate documents.

[0182] If the target binary code of any of the target character segments is different from the candidate binary code of the candidate character segments in the corresponding order in the second document, then a third comparison result is determined. The third comparison result is used to describe that the document to be processed and the second document are not similar documents.

[0183] The storage unit 1105 is specifically used to store the document to be processed in the candidate document library.

[0184] Optionally, in the above Figure 11 Based on the corresponding embodiments, in another embodiment of the document processing device provided in this application, the document processing device further includes a configuration unit 1106, a processing unit 1107, and a filtering unit 1108.

[0185] The configuration unit 1106 is specifically used to set a first identifier for the document to be processed and the first document after determining the second comparison result if the target binary code of each target character segment is the same as the candidate binary code of the candidate character segments in the corresponding order. The first identifier is used to indicate that the document to be processed and the first document are similar documents.

[0186] Processing unit 1107 is specifically used to not save the document to be processed in the candidate document library based on the first identifier;

[0187] or,

[0188] The filtering unit 1108 is specifically used to filter the document to be processed based on the first identifier.

[0189] Optionally, in the above Figure 11 Based on one or more corresponding embodiments, in another embodiment of the document processing apparatus provided in this application, the determining unit 1102 is specifically used for:

[0190] Extract the tags of each candidate domain lexicon from the multiple candidate domain lexicons. The tags of each candidate domain lexicon are used to describe the semantic domain to which the corresponding candidate domain lexicon belongs.

[0191] The document tags of the document to be processed are matched with the tags of each of the candidate domain thesauruses, so that the candidate domain thesauruses with matching tags are used as the target domain thesauruses matched by the document to be processed.

[0192] Optionally, in the above Figure 11 Based on the corresponding embodiments, in another embodiment of the document processing apparatus provided in this application, the target domain lexicon includes multiple candidate words and the word weight of each candidate word; the determining unit 1102 is specifically used for:

[0193] The document to be processed is segmented into multiple words, and each word is matched with multiple candidate words in the target domain thesaurus to determine the target candidate word that matches each word segment.

[0194] The word weight of each target candidate word is used as the word weight of the corresponding text segmentation.

[0195] Optionally, in the above Figure 11 Based on one or more corresponding embodiments, in another embodiment of the document processing apparatus provided in this application, the document processing apparatus further includes a generation unit 1109 and an assignment unit 1110. Specifically, the acquisition unit 1101 is further configured to: acquire the candidate document library, which includes multiple candidate documents, before determining the target domain thesaurus matched by the document to be processed from multiple candidate domain thesauruses based on the document tags of the document to be processed;

[0196] The determining unit 1102 is specifically used to extract the document tags of each candidate document, and the document tags of each candidate document are used to describe the semantic domain to which the corresponding candidate document belongs;

[0197] The determining unit 1102 is specifically used to determine the target candidate word in each of the candidate documents, wherein the target candidate word is a word whose frequency of occurrence in the corresponding candidate document is greater than a preset threshold;

[0198] The generation unit 1109 is specifically used to classify the target candidate words of multiple candidate documents based on the document tags of each candidate document, and generate multiple candidate domain lexicons, each of which is associated with a semantic domain;

[0199] The assignment unit 1110 is specifically used to assign weights to target candidate words in each candidate domain thesaurus based on the document influence of each target candidate word in the corresponding candidate document, so as to obtain the word weights of each target candidate word in each candidate domain thesaurus.

[0200] Optionally, in the above Figure 11 Based on one or more corresponding embodiments, in another embodiment of the document processing apparatus provided in this application, the determining unit 1102 is specifically used for:

[0201] Text features are extracted from multiple text segments of the document to be processed to obtain the text feature vector of each text segment in the document to be processed.

[0202] Based on the word weight of each text segment, the text feature vector of the corresponding text segment is weighted to obtain the weighted feature vector of each text segment;

[0203] The weighted feature vectors of all the text segmentation words in the document to be processed are merged to obtain the fusion feature of the document to be processed.

[0204] The fused features are subjected to feature dimensionality reduction processing to obtain the target hash signature of the document to be processed.

[0205] Optionally, in the above Figure 11 Based on one or more of the corresponding embodiments, in another embodiment of the document processing apparatus provided in this application, the determining unit 1102 is specifically used to: calculate the hash value of each text segment of the document to be processed, and obtain the text feature vector corresponding to the text segment.

[0206] Optionally, in the above Figure 11 Based on one or more corresponding embodiments, in another embodiment of the document processing apparatus provided in this application, the determining unit 1102 is specifically used for:

[0207] Extract the document title and text content of the document to be processed;

[0208] Based on a preset document classification model, the document title and text content of the document to be processed are classified to obtain the document tags of the document to be processed.

[0209] The document processing apparatus in the embodiments of this application has been described above from the perspective of modular functional entities. The document processing device in the embodiments of this application will be described below from the perspective of hardware processing. Figure 12This is a schematic diagram of the hardware structure of the document processing device provided in an embodiment of this application. The document processing device can vary considerably due to differences in configuration or performance, and may include, but is not limited to, [other types]. Figure 11 The document processing device shown in the figure, etc.

[0210] like Figure 12 As shown, the document processing device may include one or more central processing units (CPUs) 322 (e.g., one or more processors) and a memory 332, and one or more storage media 330 (e.g., one or more mass storage devices) for storing application programs 342 or data 344. The memory 332 and storage media 330 may be temporary or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), each module may include a series of instruction operations on the document processing device. Furthermore, the CPU 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the document processing device 300. Exemplarily, the CPU 322 is used to execute the application program 342 stored in the storage media 330, thereby implementing the document processing method provided in the above embodiments of this application.

[0211] The document processing device 300 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0212] For example, Figure 12 The central processing unit 322 can invoke computer execution instructions stored in memory 332 to cause the document processing device to perform actions such as... Figures 3 to 10 The method in the corresponding method embodiment.

[0213] Specifically, Figure 11 The functions / implementation processes of the determination unit 1102, word segmentation unit 1103, comparison unit 1104, storage unit 1105, configuration unit 1106, processing unit 1107, filtering unit 1108, generation unit 1109, and assignment unit 1110 can be understood through... Figure 12 The central processing unit 322 in the memory calls computer execution instructions stored in the memory 332 to achieve this. Figure 11 The function / implementation process of the acquisition unit 1101 can be achieved through... Figure 12 The input / output interface 358 is used to implement this.

[0214] The steps performed by the document processing device in the above embodiments can be based on this Figure 12 The document processing device structure is shown.

[0215] This application also provides a computer-readable storage medium storing a computer program or instructions thereon, which, when executed by a processor, implements the steps of the methods described in the foregoing embodiments.

[0216] This application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of the methods described in the foregoing embodiments.

[0217] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0218] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0219] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0220] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0221] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0222] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0223] A computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, they generate, in whole or in part, the processes or functions according to embodiments of this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., SSDs), etc.

[0224] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A document processing method, characterized in that, include: Obtain the document to be processed and determine the document tags of the document to be processed, wherein the document tags of the document to be processed are used to describe the semantic domain to which the document to be processed belongs; The document to be processed is segmented into words to obtain multiple text segments of the document to be processed; Based on the document tags of the document to be processed, the target domain thesaurus that the document to be processed matches is determined from multiple candidate domain thesauruses, and the word weight of each text segment of the document to be processed is determined from the target domain thesaurus. The word weight is used to describe the importance of the corresponding text segment in the semantic domain. Based on multiple text segmentations of the document to be processed and the word weight of each text segment in the document to be processed, the target hash signature of the document to be processed is determined; The target hash signature of the document to be processed is compared with the hash signatures of multiple candidate documents in the candidate document library to determine the comparison result. The comparison result is used to reflect whether the document to be processed and the candidate documents are similar documents.

2. The method according to claim 1, characterized in that, The comparison result is determined by comparing the target hash signature of the document to be processed with the hash signatures of multiple candidate documents in the candidate document library, including: The target hash signature of the document to be processed is split into multiple target character segments arranged in order, and each target character segment includes target binary code; The hash signature of each candidate document is split into strings to obtain multiple candidate character segments of each candidate document arranged in the order stated above, and each candidate character segment includes candidate binary code; The comparison result is determined based on the target binary codes of multiple target character segments in the document to be processed and the candidate binary codes of multiple candidate character segments in each candidate document.

3. The method according to claim 2, characterized in that, Based on the target binary codes of multiple target character segments in the document to be processed, and the candidate binary codes of multiple candidate character segments in each candidate document, the comparison result is determined, including: In the order described above, the target binary codes of multiple target character segments are compared with the candidate binary codes of multiple candidate character segments in the first document, wherein the first document is any one of the multiple candidate documents; If the target binary code of any of the target character segments is different from the candidate binary code of the candidate character segments in the corresponding order, then a first comparison result is determined. The first comparison result is used to describe that the document to be processed is not a similar document to the first document. or, If the target binary code of each target character segment is the same as the candidate binary code of the candidate character segments in the corresponding order, then a second comparison result is determined. The second comparison result is used to describe that the document to be processed is a similar document to the first document.

4. The method according to claim 3, characterized in that, If the target binary code of any of the target character segments is different from the candidate binary code of the candidate character segments in the corresponding order, then after determining the first comparison result, the method further includes: In the order described above, the target binary codes of multiple target character segments are compared with the candidate binary codes of multiple candidate character segments in the second document, wherein the second document is any document other than the first document among the multiple candidate documents; If the target binary code of any of the target character segments is different from the candidate binary code of the candidate character segments in the corresponding order in the second document, then a third comparison result is determined. The third comparison result is used to describe that the document to be processed and the second document are not similar documents. The document to be processed is saved in the candidate document library.

5. The method according to claim 3, characterized in that, If the target binary code of each target character segment is the same as the candidate binary code of the candidate character segments in the corresponding order, then after determining the second comparison result, the method further includes: A first identifier is set for the document to be processed and the first document, the first identifier being used to indicate that the document to be processed and the first document are similar documents; Based on the first identifier, the document to be processed is not saved in the candidate document library; or, The documents to be processed are filtered based on the first identifier.

6. The method according to any one of claims 1 to 5, characterized in that, Based on the document tags of the document to be processed, the target domain thesaurus that matches the document to be processed is determined from multiple candidate domain thesauruses, including: Extract the tags of each candidate domain lexicon from the multiple candidate domain lexicons. The tags of each candidate domain lexicon are used to describe the semantic domain to which the corresponding candidate domain lexicon belongs. The document tags of the document to be processed are matched with the tags of each of the candidate domain thesauruses, so that the candidate domain thesauruses with matching tags are used as the target domain thesauruses matched by the document to be processed.

7. The method according to any one of claims 1 to 5, characterized in that, The target domain lexicon includes multiple candidate words and the word weight of each candidate word; Determining the word weight of each text segment of the document to be processed from the target domain lexicon includes: The document to be processed is segmented into multiple words, and each word is matched with multiple candidate words in the target domain thesaurus to determine the target candidate word that matches each word segment. The word weight of each target candidate word is used as the word weight of the corresponding text segmentation.

8. The method according to any one of claims 1 to 7, characterized in that, Before determining the target domain thesaurus that matches the document to be processed from multiple candidate domain thesauruses based on the document tags of the document to be processed, the method further includes: Obtain the candidate document library, which includes multiple candidate documents; Extract the document tag for each candidate document, and the document tag for each candidate document is used to describe the semantic domain to which the corresponding candidate document belongs; Identify target candidate words in each candidate document, wherein the target candidate words are words that appear more frequently than a preset threshold in the corresponding candidate document; Based on the document tags of each candidate document, the target candidate words of multiple candidate documents are classified to generate multiple candidate domain lexicons, and each candidate domain lexicon is associated with a semantic domain; Based on the document influence of each target candidate word in the corresponding candidate document, weights are assigned to each target candidate word in each candidate domain thesaurus to obtain the word weights of each target candidate word in each candidate domain thesaurus.

9. The method according to any one of claims 1 to 8, characterized in that, Based on multiple text segmentations of the document to be processed and the word weight of each text segment in the document to be processed, the target hash signature of the document to be processed is determined, including: Text features are extracted from multiple text segments of the document to be processed to obtain the text feature vector of each text segment in the document to be processed. Based on the word weight of each text segment, the text feature vector of the corresponding text segment is weighted to obtain the weighted feature vector of each text segment; The weighted feature vectors of all the text segmentation words in the document to be processed are merged to obtain the fusion feature of the document to be processed. The fused features are subjected to feature dimensionality reduction processing to obtain the target hash signature of the document to be processed.

10. The method according to claim 9, characterized in that, Text features are extracted from multiple text segments of the document to be processed to obtain a text feature vector for each of the text segments in the document to be processed, including: Calculate the hash value of each text segment of the document to be processed to obtain the text feature vector corresponding to the text segment.

11. The method according to any one of claims 1 to 10, characterized in that, Determining the document tags of the document to be processed includes: Extract the document title and text content of the document to be processed; Based on a preset document classification model, the document title and text content of the document to be processed are classified to obtain the document tags of the document to be processed.

12. A document processing apparatus, characterized in that, include: The acquisition unit is used to acquire the document to be processed. A determining unit is used to determine the document tags of the document to be processed, wherein the document tags of the document to be processed are used to describe the semantic domain to which the document to be processed belongs; The word segmentation unit is used to segment the document to be processed into multiple text words of the document to be processed. The determining unit is configured to determine the target domain thesaurus that matches the document to be processed from multiple candidate domain thesauruses based on the document tags of the document to be processed, and to determine the word weight of each text segment of the document to be processed from the target domain thesaurus, wherein the word weight is used to describe the importance of the corresponding text segment in the semantic domain; The determining unit is used to determine the target hash signature of the document to be processed based on multiple text segments of the document to be processed and the word weight of each text segment in the document to be processed; The comparison unit is used to compare the target hash signature of the document to be processed with the hash signatures of multiple candidate documents in the candidate document library, and determine the comparison result. The comparison result is used to reflect whether the document to be processed and the candidate documents are similar documents.

13. A document processing device, characterized in that, include: Input / output interface, processor, and memory, wherein the memory stores program instructions; The processor is configured to execute program instructions stored in the memory to perform the method as described in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed on a computer device, cause the computer device to perform the method as described in any one of claims 1 to 11.

15. A computer program product, characterized in that, The computer program product includes instructions that, when executed on a computer device, cause the computer device to perform the method as described in any one of claims 1 to 11.