Text corpus processing method and apparatus, storage medium, and electronic device
By segmenting and extracting features from the text corpus, and improving the hash algorithm by combining semantics and positional importance, the problem of low efficiency in text corpus deduplication is solved, achieving efficient and accurate text corpus deduplication, which is suitable for distributed storage scenarios.
Patent Information
- Application Number
- CN202111415376.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2041-11-25
AI Technical Summary
Existing technologies for deduplication of text corpora are inefficient and inaccurate, making them unsuitable for deduplication scenarios involving massive amounts of text corpora, and they also have poor compatibility in distributed storage scenarios.
By segmenting the text corpus, extracting the feature and weight information of the word sequence, and combining semantic importance and positional importance, improve the hash algorithm to perform hash mapping and weighted encoding, and build an index to remove duplicates.
It significantly improves the speed and accuracy of text corpus deduplication, making it suitable for distributed storage scenarios with massive amounts of data and reducing the pressure of redundant text on downstream analysis.
Smart Images

Figure CN114328818B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a text corpus processing method and device, a storage medium and an electronic device. BACKGROUND
[0002] With the development of computer technology, applications relying on text information analysis have been increasingly popular, such as advertisement recommendation, news promotion, sharing of various media content, and the like, which all rely on analysis of text information. In order to reduce the pressure of text information analysis, it is necessary to perform deduplication operation on massive text corpus. In related technologies, the text information extraction accuracy of text corpus is low and slow, resulting in low efficiency and poor effect of deduplication operation of text corpus, and it is also difficult to be applied to the deduplication scene of massive text corpus. SUMMARY
[0003] To solve at least one of the above technical problems, embodiments of the present application provide a text corpus processing method and device, a storage medium and an electronic device.
[0004] In one aspect, a text corpus processing method is provided, which includes:
[0005] obtaining a text corpus;
[0006] performing word segmentation processing on the text corpus to obtain a word sequence corresponding to the text corpus;
[0007] performing information extraction processing on the word sequence to obtain feature information and weight information corresponding to each word in the word sequence, the weight information being determined according to semantic importance and position importance of the word in the word sequence;
[0008] performing hash mapping on the feature information corresponding to each word to obtain encoding information corresponding to each word;
[0009] obtaining weighted encoding information corresponding to each word according to the encoding information corresponding to each word and the corresponding weight information;
[0010] performing fusion operation on the encoding information corresponding to each word to obtain text information corresponding to the text corpus.
[0011] In another aspect, a text corpus processing device is provided, which includes:
[0012] a text corpus obtaining module configured to obtain a text corpus;
[0013] an analysis module configured to perform word segmentation processing on the text corpus to obtain a word sequence corresponding to the text corpus;
[0014] The information extraction module is configured to perform information extraction processing on the word sequence to obtain feature information and weight information corresponding to each word in the word sequence, wherein the weight information is determined according to semantic importance and position importance of the word in the word sequence.
[0015] The hashing module is configured to perform hash mapping on the feature information corresponding to each word to obtain encoding information corresponding to each word.
[0016] The weighting module is configured to obtain weighted encoding information corresponding to each word according to the encoding information corresponding to each word and the weight information corresponding to each word.
[0017] The fusion module is configured to perform fusion operation on the encoding information corresponding to each word to obtain text information corresponding to the text corpus.
[0018] In another aspect, an embodiment of the present application provides a distributed storage system, which comprises the text corpus processing apparatus described above.
[0019] In another aspect, an embodiment of the present application provides a computer readable storage medium, which stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the text corpus processing method described above.
[0020] In another aspect, an embodiment of the present application provides an electronic device, which comprises at least one processor and a memory connected to the at least one processor in communication, and the memory stores instructions executable by the at least one processor, and the at least one processor implements the text corpus processing method described above by executing the instructions stored in the memory.
[0021] In another aspect, an embodiment of the present application provides a computer program product, which comprises a computer program or instructions, and the computer program or instructions are executed by a processor to implement the text corpus processing method described above.
[0022] The text corpus processing method provided by the present application improves the hash algorithm based on the features of the words in the text corpus, in combination with the semantic importance and position importance of the text corpus, so as to obtain text information that can more accurately represent the information in the text corpus, and the text corpus deduplication based on the text information can significantly improve the text corpus deduplication speed and accuracy. Moreover, the text information extraction method is applied to the scenario of massive data distributed storage, so that the text information in the scenario of massive data distributed storage can be extracted in time, and the text corpus can be deduplicated in time, thereby avoiding the pressure on the downstream text analysis caused by redundant text corpora. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the accompanying drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on these drawings.
[0024] Figure 1 is a feasible implementation framework schematic diagram of the text corpus processing method provided by the embodiments of the present application;
[0025] Figure 2 is a flowchart schematic diagram of the text corpus processing method provided by the embodiments of the present application;
[0026] Figure 3 is a flowchart of the distributed system deduplication method provided by the embodiments of the present application;
[0027] Figure 4 is a flowchart schematic diagram of the first identification or second identification based query method provided by the embodiments of the present application;
[0028] Figure 5 is a block diagram of the text corpus processing apparatus provided by the embodiments of the present application;
[0029] Figure 6 is a hardware structure schematic diagram of the device for implementing the method provided by the embodiments of the present application. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the present application will be described clearly and completely with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above drawings are used for distinguishing similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged, where appropriate, so that the embodiments of the present application described herein can be carried out in other than the order shown or described herein. Furthermore, the terms "comprising" and "having", and any variations thereof, are intended to cover a non-exclusive inclusion, for example, a process, method, system, product or server that comprises a list of steps or units is not necessarily limited to those steps or units that are clearly listed, but can include other steps or units that are not clearly listed or inherent to such processes, methods, products or devices.
[0032] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application are further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, and do not limit the embodiments of the present application.
[0033] Hereinafter, the terms "first" and "second" are only used for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more features. In the description of the embodiments, unless otherwise specified, the meaning of "multiple" is two or more. In order to facilitate understanding of the above technical solutions of the embodiments of the present application and the technical effects generated thereby, the embodiments of the present application first explain the related professional terms:
[0034] Artificial Intelligence (AI): is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0035] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0036] Convolutional Neural Networks, CNN. is a kind of feedforward neural network containing convolution calculation and having deep structure, which is one of the representative algorithms of deep learning. Convolutional neural network has the ability of representation learning and can perform translation-invariant classification on input information according to its hierarchical structure, so it is also called translation-invariant artificial neural network.
[0037] Recurrent Neural Network, RNN. is a kind of recurrent neural network with sequence data as input, which performs recursion in the evolution direction of sequence and all nodes (recurrent units) are connected in chain.
[0038] Bidirectional Encoder Representation from Transformers, BERT. is a model for pre-training language representation, which trains a general "language understanding" model based on text corpus. Based on BERT model, natural language processing (NLP) tasks can be assisted.
[0039] Spark. is a fast general-purpose computing engine designed for big data processing. Spark is an open source cluster computing environment that enables in-memory distributed datasets. In addition to providing interactive queries, it can also optimize iterative workloads and easily manipulate distributed datasets like local collection objects.
[0040] HDFS. Hadoop Distributed File System, distributed file system. HDFS is a highly fault-tolerant system suitable for deployment on inexpensive machines. HDFS can provide high-throughput data access and is very suitable for large-scale data sets.
[0041] Redis. Remote Dictionary Server, remote dictionary service, is an open source network-enabled, in-memory-based, log-based, key-value database that can also be persistent.
[0042] BitMap. Use a bit to mark the value corresponding to an element, and the key is the element. Because the data is stored in bits, the storage space can be greatly saved.
[0043] Hamming distance: in an effective bit encoding set, XOR operation is performed on two bit strings, and the number of 1s in the XOR operation result is the Hamming distance of the two bit strings, or the Hamming distance of two strings of equal length is the number of different characters in the same position.
[0044] SimHash algorithm: the traditional hash algorithm may have a huge difference in the hash value generated if two original contents only differ by a few bytes. If the original contents are not much different, the SimHash algorithm will also have a small difference in the hash value.
[0045] The related technology for deduplication of text corpus can rely on the Hash method, and the traditional Hash method is only responsible for mapping the original content as evenly and randomly as possible to a signature value. Even if two texts only differ by a few characters, the hash value obtained may also differ greatly. Therefore, the traditional hash algorithm is difficult to measure the similarity of the contents. The SimHash algorithm is a data locality sensitive hash, and its main idea is to reduce the dimension, convert the high-dimensional feature vector into a fixed length hash value, and determine the similarity of two texts by calculating the distance between two hash values. However, the SimHash algorithm in the related technology has less consideration for the semantic importance and position importance of the words in the text corpus, which to some extent affects the deduplication accuracy, and the combination of cosine similarity to determine whether the text corpus is repeated or similar also reduces the deduplication speed to some extent.
[0046] In addition, the deduplication operation of the related technology for the text corpus does not consider the storage scenario of massive text corpus in a distributed scenario, and the text corpus deduplication method is difficult to effectively compatible and adapt to the storage structure of the text corpus, which also affects the deduplication efficiency of the text corpus to some extent.
[0047] In order to improve the deduplication accuracy and speed of the text corpus, so that the text corpus can be more efficiently deduplicated in the massive data distributed storage scenario, and reduce the pressure of analysis based on the text corpus, the embodiments of the present application provide a text corpus processing method. Based on the features of the words in the text corpus, combined with the semantic importance and position importance of the text corpus, the hash algorithm is improved, so that the text information which can more accurately represent the information in the text corpus is obtained, and the deduplication of the text corpus based on the text information can significantly improve the deduplication speed and accuracy of the text corpus. In addition, the text information extraction method is also applied to the massive data distributed storage scenario, so that the text information in the massive data distributed storage scenario can be extracted in time, and the text corpus can be deduplicated in time, thereby avoiding the pressure of redundant text corpus on downstream text analysis.
[0048] The embodiments of the present application can relate to cloud technology and cloud gaming. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, network, etc. in a wide area network or local area network to realize data calculation, storage, processing and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on cloud computing business model application, which can form a resource pool, and can be used on demand, flexibly and conveniently. Cloud computing technology will become an important support. The background service of a technical network system needs a large amount of computing and storage resources, such as video websites, picture websites and more portals. With the high development and application of the Internet industry, in the future, every item may have its own identification mark and needs to be transmitted to the background system for logical processing. Different levels of data will be processed separately, and various industry data need strong system support, which can only be realized through cloud computing.
[0049] The method provided by the embodiments of the present application can also relate to a blockchain, that is, the method provided by the embodiments of the present application can be implemented based on a blockchain, or the data involved in the method provided by the embodiments of the present application can be stored based on a blockchain, or the execution subject of the method provided by the embodiments of the present application can be located in a blockchain. The blockchain is a new application mode of distributed data storage, point-to-point transmission, consensus mechanism, encryption algorithm and other computer technologies. Blockchain, in essence, is a decentralized database, which is a series of data blocks associated using cryptography. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-fake) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer and an application service layer.
[0050] The blockchain underlying platform can include user management, basic services, smart contracts, and operation monitoring processing modules. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining public and private key generation (account management), key management, and user real identity and blockchain address correspondence maintenance (permission management), etc., and under authorization, supervising and auditing the transaction of certain real identities, providing risk control rule configuration (risk audit); the basic service module is deployed on all blockchain node devices to verify the validity of business requests, and record to the storage after consensus for valid requests, for a new business request, the basic service first interface adaptation analysis and authentication processing (interface adaptation), then encrypt the business information through the consensus algorithm (consensus management), after encryption, the complete and consistent transmission to the shared ledger (network communication), and record storage; the smart contract module is responsible for contract registration and issuance, contract triggering and contract execution, developers can define contract logic through a certain programming language, publish to the blockchain (contract registration), according to the logic of the contract terms, call the key or other event triggers to execute, complete the contract logic, and also provide contract upgrade and cancellation functions; the operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation in the product release process, and the real-time state visualization output in the product running, such as: alarm, monitoring network situation, monitoring node device health status, etc.
[0051] The platform product service layer provides basic capabilities and implementation framework of typical applications, and developers can stack business characteristics based on the basic capabilities to complete the blockchain implementation of business logic. The application service layer provides application services based on the blockchain scheme for business participants to use.
[0052] Please refer to Figure 1 , Figure 1 is a feasible implementation framework diagram of the text corpus processing method provided by the embodiment of the present specification, as Figure 1 shown, the implementation framework can run a distributed storage system, specifically at least including a data storage component 01, a data management component 02, and a data analysis component 03. Among them, the data storage component 01, the data management component 02, and the data analysis component 03 can be devices located in the Internet, and can provide various optional Internet-based services for users, including but not limited to mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, independent servers or servers running in a cloud environment, etc.
[0053] The data storage component 01 in the embodiment of the present application can run HDFS, the data management component 02 can run Redis, and the data analysis component 03 can run Spark. The data storage component 01 stores all data of the deduplicated text corpus, the data management component 02 stores the first index obtained based on the hash value corresponding to the title of the deduplicated text corpus, the second index obtained based on the hash value corresponding to the content corpus of the deduplicated text corpus, and the text information of the deduplicated text corpus. The data analysis component 03 implements text deduplication by interacting with the data management component 02.
[0054] The following describes a text corpus processing method according to an embodiment of the present application, Figure 2 A flowchart of a text corpus processing method according to an embodiment of the present application is shown. The embodiment of the present application provides the method operation steps as described above in the embodiment or flowchart, but more or fewer operation steps can be included based on conventional or non-inventive labor. The order of steps listed in the embodiment is only one of the many execution orders, and does not represent the only execution order. In actual system, terminal device or server product execution, the method order shown in the embodiment or the drawing can be executed in sequence or in parallel (for example, parallel processor or multi-threaded processing environment), and the above method can include:
[0055] S101. Obtain a text corpus.
[0056] The embodiment of the present application does not limit the type of text corpus. For example, the text corpus can be information stored in various text forms from the Internet, such as from community websites, instant messaging records, media content sharing platforms, etc. The text corpus in the embodiment of the present application can be corpus ready to be stored in a distributed storage system. Before storage, the method of the present application is executed to check whether the text corpus is duplicated or similar to the existing text corpus in the distributed storage system. If so, the text corpus is discarded directly, otherwise, the text corpus is stored.
[0057] S102. Perform word segmentation processing on the text corpus to obtain a word sequence corresponding to the text corpus.
[0058] The text corpus in the embodiment of the present application can include two parts of content, i.e. title and content, that is, the text corpus includes title corpus and content corpus. Taking a Chinese form text corpus as an example, the text corpus "today the weather is really good" can be segmented into a word sequence including three words {‘today’, ‘weather’, ‘really good’}. Taking an English form text corpus as an example, the text corpus "david is a cute boy" can be segmented into a word sequence including five words {‘david’, ‘is’, ‘a’, ‘cute’, ‘boy’}.
[0059] The embodiment of the present application does not limit the word segmentation method used. For example, a word segmentation method based on string matching can be used, which means that the Chinese character string to be analyzed is matched with the entries in a pre-set machine dictionary according to a certain strategy, and word segmentation is performed according to the matching result. A word segmentation method based on feature scanning or marker segmentation can also be used, which preferentially identifies and segments some words with obvious features in the string to be analyzed, uses these words as breakpoints, and divides the original string into smaller strings for mechanical word segmentation, thereby improving the accuracy of word segmentation. A word segmentation method based on understanding can also be used, which means that the computer simulates human understanding of a sentence to achieve the effect of identifying words. The basic idea is to perform syntax and semantic analysis at the same time of word segmentation, and use syntax information and semantic information to handle ambiguity. Or a word segmentation method based on statistics can also be used, since the frequency or probability of the co-occurrence of adjacent words can better reflect the credibility of word formation, the frequency of the combination of adjacent co-occurrence words in the corpus can be counted, the mutual information of the words can be calculated, and word segmentation can be performed based on the mutual information.
[0060] S103. Information extraction processing is performed on the above word sequence to obtain feature information and weight information corresponding to each word in the above word sequence, wherein the weight information is determined according to the semantic importance and position importance of the word in the above word sequence.
[0061] The embodiment of the present application does not limit the specific operation method of the information extraction processing. For example, it can be implemented by relying on a neural network, such as a supervised neural network model or an unsupervised neural network model, and the BERT model is described in detail.
[0062] Specifically, a sample corpus can be obtained, the title and text of the sample corpus are subjected to word segmentation processing to obtain a sample word sequence, and a human being annotates the weight based on the semantic importance and position importance of each word in the sample word sequence to obtain an annotation result. In an embodiment, the human being can score the word according to the semantic importance, and the score is between 0 and 1, and the higher the score, the more important the semantic of the word is to the understanding of the entire sample corpus. The position of the word in the corpus is determined, the position importance of the word is determined according to the position, the position importance can also be represented by a number between 0 and 1, and the semantic importance and position importance are weighted and averaged to obtain the annotation result. Of course, the weight is not limited in the embodiment of the present application, and can be set according to the actual situation. The BERT model is trained based on the sample corpus carrying the annotation result, so that the model can predict the weight information of each word in the corpus, and the BERT model parameters can be corrected according to the prediction result and the annotation result. The word sequence in step S102 is input into the trained BERT model, and the feature information and weight information corresponding to each word in the above word sequence are obtained.
[0063] S104. Hash mapping the feature information corresponding to each of the above words to obtain the encoding information corresponding to each of the above words.
[0064] Specifically, the feature vector corresponding to each of the above words can be scattered by hash to obtain the first encoding string corresponding to each of the above words, and the first encoding string is a data string composed of binary numbers. The 0 value in the first encoding string is set as a preset negative value to obtain the encoding information corresponding to each of the above words. The preset negative value can be any integer with a negative value set according to the actual situation, such as -1, -2, etc.
[0065] Taking the word "we" in the word sequence as an example, the first encoding string corresponding thereto can be obtained by hash scattering, and the first encoding string is a data string composed of binary numbers 0 or 1, which can be stored in the form of BitMap. For example, the first encoding string corresponding to the word "we" is [1, 1, 0, 1, 0, 1]. On the basis of the first encoding string, all 0 values of the bit positions in the first encoding string are converted to -1. For example, the first encoding string corresponding to "we" is [1, 1, 0, 1, 0, 1, and the converted encoding information is [1, 1, -1, 1, -1, 1].
[0066] S105. Obtaining the weighted encoding information corresponding to each of the above words according to the encoding information corresponding to each of the above words and the corresponding weight information.
[0067] Specifically, the weighted encoding information can be directly determined according to the product of the encoding information corresponding to each of the above words and the corresponding weight information. Still taking "we" as an example, if the weight value corresponding thereto is 3, the weighted encoding information corresponding thereto is [3, 3, -3, 3, -3, 3].
[0068] S106. Performing a fusion operation on the encoding information corresponding to each word to obtain the text information corresponding to the above text corpus.
[0069] Specifically, the encoding information corresponding to each word can be added bit by bit to obtain the fusion encoding information. The fusion encoding information is subjected to a dimension reduction operation to obtain the text information corresponding to the above text corpus.
[0070] Specifically, the weighted encoding information of all words can be accumulated, and the positions greater than 0 in the accumulation result are 1, and the positions less than 0 are 0, thereby obtaining the text information corresponding to the text corpus.
[0071] To illustrate this more clearly, let's take another example. The first encoded string for "life" is "110101," and the weighted encoded information can be "5, 5, -5, 5, -5, 5." The first encoded string for "nothing" is "101001," and the weighted encoded information can be "2, 2, 2, -2, -2, 2." The text information corresponding to the corpus "life nothing" is obtained as follows: "5, 5, -5, 5, -5, 5" and "2, 2, 2, -2, -2, 2" are summed to obtain "7, 3, -3, 3, -7, 7." If the sum is greater than 0, it is set to 1; otherwise, it is set to 0, resulting in the final text information "1, 1, 0, 1, 0, 1."
[0072] In one embodiment, such as Figure 3 As shown, the above method can be applied to a distributed storage system, which includes a data analysis component, a data storage component, and a data management component. Before performing word segmentation on the text corpus to obtain the word sequence corresponding to the text corpus, the method further includes:
[0073] S201. The above data analysis component extracts the title corpus and content corpus from the above text corpus.
[0074] This application does not limit the extraction method. Data analysis components can be used to extract title corpora and content corpora. If the data analysis component includes Spark, Spark's inherent operators can be used to execute step S201.
[0075] S202. In response to the first case, which indicates that the title corpus and the content corpus are not duplicates of the corpus already stored in the distributed storage system, the data analysis component performs the operation of segmenting the text corpus to obtain the word sequence corresponding to the text corpus.
[0076] If the title corpus overlaps with the title corpus of an existing text corpus in the distributed storage system, a case of title corpus duplication occurs. Similarly, if the content corpus overlaps with the content corpus of an existing text corpus in the distributed storage system, a case of content corpus duplication occurs. In the distributed storage system, after obtaining the text corpus, step S101 is executed only if neither of these two types of duplication exists (the first case). Otherwise, the corpus is discarded to avoid storing already stored text corpus in the distributed storage system.
[0077] S203. In response to other situations different from the first situation, the above data analysis component discards the above text corpus.
[0078] In an embodiment, before the operation of performing the above-described word segmentation processing on the above-described text corpus to obtain a word sequence corresponding to the text corpus, the method further includes, in response to the first case, performing the operation of:
[0079] S301. The data analysis component obtains a first identifier or a second identifier, the first identifier being a hash value corresponding to the title corpus, and the second identifier being a hash value corresponding to the content corpus.
[0080] S302. The data analysis component accesses the data management component based on the first identifier or the second identifier to obtain a query result fed back by the data management component, the query result representing a duplication situation of the text corpus and the content corpus.
[0081] Embodiments of the present application do not limit the method of obtaining the query result, and a structured query language can be used for querying, for example, querying the data management component for records that hit the first identifier or records that hit the second identifier. Through interaction with the data management component, the existing corpus that is the same as the text corpus title or the same as the content in step S101 can be queried out. If the queried result is not empty, it indicates that there is a duplication situation, and the text corpus in step S101 can be abandoned to avoid duplication storage.
[0082] Specifically, referring to Figure 4 The data analysis component accesses the data management component based on the first identifier or the second identifier to obtain a query result fed back by the data management component, including:
[0083] S3021. The data analysis component sends a first query instruction to the data management component based on the first identifier.
[0084] The first query instruction is generated according to the first identifier, and is used to query the data management component for related records including the first identifier, that is, to query the existing corpus that is the same as the title of the text corpus in step S101.
[0085] S3022. In response to the first query instruction, the data management component queries a first index based on the first identifier to feed back a first query result, the first index being an index constructed based on hash values corresponding to title corpora of existing text corpora in the data storage component.
[0086] S3023. The data analysis component determines that the title corpus is duplicated in response to the first query result being not empty.
[0087] The first index is constructed based on hash values corresponding to title corpora of existing corpora, and through index querying, the existence of the existing corpus that is the same as the title of the text corpus in step S101 can be quickly determined.
[0088] S3024. The data analysis component responds to the case that the first query result is empty, and sends a second query instruction to the data management component based on the second identifier.
[0089] The second query instruction is generated according to the second identifier, and is used to query the data management component for a related record including the second identifier, that is, to query an existing corpus that has the same content as the text corpus in step S101.
[0090] S3025. In response to the second query instruction, the data management component queries the second index based on the second identifier, feeds back a second query result, and the second index is an index constructed based on the hash value corresponding to the content corpus of the existing text corpus in the data storage component.
[0091] S3026. The data analysis component responds to the case that the second query result is not empty, and determines that the content corpus is duplicated.
[0092] The second index is constructed based on the hash value corresponding to the content corpus in the existing corpus, and through index query, the existence of the existing corpus having the same content as the text corpus in step S101 can be quickly determined.
[0093] S3027. The data analysis component responds to the case that the second query result is empty, and determines that neither the title corpus nor the content corpus is duplicated.
[0094] The embodiments of the present application realize the quick query of the title duplication case and the quick query of the content duplication case by constructing the first index based on the hash value of the title and the second index based on the hash value of the content, and can significantly improve the text corpus duplication checking speed.
[0095] In one embodiment, in response to the first case, the method further comprises:
[0096] S401. The data analysis component sends a third query instruction to the data management component based on the text information corresponding to the text corpus, so as to obtain a third query result fed back by the data management component, and the third query result represents whether there is similar corpus corresponding to the text corpus in the data storage component.
[0097] Specifically, the data management component obtains a text information table, where each record represents the first preset number of digits of information corresponding to the first text information already stored in the data storage component. Based on the first preset number of digits of information in the text information, the data management component queries the text information table to obtain a third query result. The similarity between the records in the third query result and the first preset number of digits of information in the text information meets a preset requirement.
[0098] For each text corpus i stored in the data storage component, the embodiments of this application can refer to the operations of steps S101-S106 above to extract the first text information corresponding to the text corpus i. This first text information is obtained through V. i This means that the V can be... i Divide into n parts, where V i The length is m, and each piece has m / n characters. The value of each piece is stored in the data management component. Based on the first m / n characters of each first text information, a text information table can be obtained. To improve query speed, an exact match can be used to find the first m / n records. If a record that meets the requirements is found, the third query result mentioned above is obtained based on that record.
[0099] To improve query speed, this application embodiment uses Hamming distance. Specifically, the data management component calculates the Hamming distance between the first preset number of digits of the text information and each record in the text information table; based on the records whose Hamming distances meet preset requirements, the third query result is obtained.
[0100] Specifically, the smaller the Hamming distance, the lower the similarity. It is generally considered that a Hamming distance of 3 represents that the two articles are identical. Therefore, if there is a record with a Hamming distance less than 3, it can be considered that the record is similar to the text in the corpus in step S101.
[0101] S402. In response to the third query result indicating that no similar corpus exists, the text corpus is stored in the data storage component, and the text information corresponding to the text corpus, the first identifier, and the second identifier are stored in the data management component.
[0102] S403. In response to the above third query result indicating the existence of similar corpus, discard the above text corpus.
[0103] In order to improve the deduplication accuracy and speed of the text corpus, so that the text corpus can be more efficiently deduplicated in the massive data distributed storage scene, and reduce the pressure of analysis based on the text corpus, an embodiment of the present application provides a text corpus processing method. Based on the features of the words in the text corpus, combined with the semantic importance and position importance of the text corpus, the hash algorithm is improved, so that the text information which can more accurately represent the information in the text corpus is obtained. Based on the text information, the deduplication of the text corpus can significantly improve the deduplication speed and accuracy of the text corpus. Moreover, the text information extraction method is applied to the massive data distributed storage scene, so that the text information in the massive data distributed storage scene can be extracted in time, and the text corpus can be deduplicated in time, thereby avoiding the pressure of redundant text corpus on downstream text analysis.
[0104] Please refer to Figure 5 which shows a block diagram of a text corpus processing device in an embodiment, the device includes:
[0105] The text corpus acquisition module 101 is configured to acquire a text corpus.
[0106] The analysis module 102 is configured to perform word segmentation processing on the text corpus to obtain a word sequence corresponding to the text corpus.
[0107] The information extraction module 103 is configured to perform information extraction processing on the word sequence to obtain feature information and weight information corresponding to each word in the word sequence, wherein the weight information is determined according to the semantic importance and position importance of the word in the word sequence.
[0108] The hash module 104 is configured to perform hash mapping on the feature information corresponding to each word to obtain encoding information corresponding to each word.
[0109] The weighting module 105 is configured to obtain weighted encoding information corresponding to each word according to the encoding information corresponding to each word and the corresponding weight information.
[0110] The fusion module 106 is configured to perform fusion operation on the encoding information corresponding to each word to obtain text information corresponding to the text corpus.
[0111] In one embodiment, the hash module 104 is configured to perform the following operations:
[0112] Hash scattering is performed on the feature vector corresponding to each word to obtain a first encoding string corresponding to each word, wherein the first encoding string is a data string composed of binary numbers.
[0113] The 0 value in the first code string is set as a preset negative value, to obtain the code information corresponding to each word.
[0114] In one embodiment, the fusion module 106 is configured to perform the following operation: performing a bitwise addition operation on the code information corresponding to each word to obtain fusion code information.
[0115] Performing a dimension reduction operation on the fusion code information to obtain the text information corresponding to the text corpus.
[0116] In one embodiment, the device is applied to a distributed storage system, and the distributed storage system includes a data analysis component. The distributed storage system is configured to perform the following operation:
[0117] The data analysis component extracts the title corpus and the content corpus in the text corpus.
[0118] In response to a first case, the data analysis component performs the operation of performing the word segmentation processing on the text corpus to obtain the word sequence corresponding to the text corpus, where the first case represents that the title corpus and the content corpus are not repeated with the existing corpus of the distributed storage system.
[0119] In response to a case different from the first case, the data analysis component discards the text corpus.
[0120] In one embodiment, the distributed storage system is configured to perform the following operation:
[0121] The data analysis component obtains a first identifier or a second identifier, where the first identifier is a hash value corresponding to the title corpus, and the second identifier is a hash value corresponding to the content corpus.
[0122] The data analysis component accesses the data management component based on the first identifier or the second identifier to obtain a query result fed back by the data management component, where the query result represents a repetition condition of the text corpus and the content corpus.
[0123] In one embodiment, the distributed storage system is configured to perform the following operation:
[0124] The data analysis component sends a first query instruction to the data management component based on the first identifier.
[0125] In response to the first query instruction, the data management component queries a first index based on the first identifier to feed back a first query result, where the first index is an index constructed based on the hash value of the title corpus of the existing text corpus in the data storage component.
[0126] The data analysis component determines that the title corpus is duplicated in response to the first query result being not empty.
[0127] The data analysis component sends a second query instruction to the data management component based on the second identifier in response to the first query result being empty.
[0128] The data management component queries a second index based on the second identifier in response to the second query instruction and feeds back a second query result, the second index being an index constructed based on hash values corresponding to content corpora of the existing text corpora in the data storage component.
[0129] The data analysis component determines that the content corpus is duplicated in response to the second query result being not empty.
[0130] The data analysis component determines that neither the title corpus nor the content corpus is duplicated in response to the second query result being empty.
[0131] In one embodiment, the distributed storage system is configured to perform the following operations:
[0132] The data analysis component sends a third query instruction to the data management component based on the text information corresponding to the text corpus, so as to obtain a third query result fed back by the data management component, the third query result representing whether similar corpus corresponding to the text corpus exists in the data storage component.
[0133] In response to the third query result representing that the similar corpus does not exist, the text corpus is stored in the data storage component, and the text information corresponding to the text corpus, the first identifier and the second identifier are stored in the data management component.
[0134] In response to the third query result representing that the similar corpus exists, the text corpus is discarded.
[0135] In one embodiment, the distributed storage system is configured to perform the following operations:
[0136] The data management component acquires a text information table, records in the text information table representing the first text information corresponding to the text corpus stored in the data storage component.
[0137] The data management component queries the text information table based on the first text information, so as to obtain a third query result, a record in the third query result satisfying a preset requirement in terms of similarity with the first text information.
[0138] In one embodiment, the distributed storage system described above is used to perform the following operations:
[0139] The data management component calculates the Hamming distance between the first preset number of bits of the text information and each record in the text information table;
[0140] According to the records whose Hamming distance meets the preset requirement, the third query result is obtained.
[0141] The device embodiment and the method embodiment of the embodiments of the present application are based on the same inventive concept, which will not be repeated here.
[0142] The embodiments of the present application also provide a distributed storage system, which comprises the text corpus processing device described above.
[0143] The embodiments of the present application also provide a computer program product or a computer program, which comprises computer instructions stored in a computer readable storage medium. The processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the text corpus processing method described above.
[0144] The embodiments of the present application also provide a computer readable storage medium, which can store a plurality of instructions. The instructions can be adapted to be loaded by a processor and execute the text corpus processing method described above.
[0145] In one embodiment, the text corpus processing method comprises:
[0146] Obtaining a text corpus;
[0147] Performing a word segmentation processing on the text corpus to obtain a word sequence corresponding to the text corpus;
[0148] Performing an information extraction processing on the word sequence to obtain feature information and weight information corresponding to each word in the word sequence, wherein the weight information is determined according to the semantic importance and the position importance of the word in the word sequence;
[0149] Hash mapping the feature information corresponding to each word to obtain encoding information corresponding to each word;
[0150] Obtaining weighted encoding information corresponding to each word according to the encoding information corresponding to each word and the corresponding weight information;
[0151] Performing a fusion operation on the encoding information corresponding to each word to obtain text information corresponding to the text corpus.
[0152] In an embodiment, the feature information is represented by a feature vector, the feature information corresponding to each word is mapped by a hash function to obtain the encoding information corresponding to each word, including:
[0153] The feature vector corresponding to each word is scattered by a hash function to obtain a first encoding string corresponding to each word, the first encoding string being a data string composed of binary numbers;
[0154] The value of 0 in the first encoding string is set to a preset negative value to obtain the encoding information corresponding to each word.
[0155] In an embodiment, the encoding information corresponding to each word is fused to obtain the text information corresponding to the text corpus, including:
[0156] The encoding information corresponding to each word is added bit by bit to obtain fused encoding information;
[0157] The fused encoding information is reduced in dimension to obtain the text information corresponding to the text corpus.
[0158] In an embodiment, the method is applied to a distributed storage system, the distributed storage system including a data analysis component, and before the text corpus is segmented to obtain the word sequence corresponding to the text corpus, the method further includes:
[0159] The data analysis component extracts the title corpus and the content corpus in the text corpus;
[0160] In response to a first condition, the first condition indicating that the title corpus and the content corpus are both not repeated with the stored corpus of the distributed storage system, the data analysis component performs the operation of segmenting the text corpus to obtain the word sequence corresponding to the text corpus;
[0161] In response to a condition different from the first condition, the data analysis component discards the text corpus.
[0162] In an embodiment, the distributed storage system further includes a data storage component and a data management component, and before the operation of segmenting the text corpus to obtain the word sequence corresponding to the text corpus, the method further includes:
[0163] The data analysis component obtains a first identifier or a second identifier, the first identifier being a hash value corresponding to the title corpus, and the second identifier being a hash value corresponding to the content corpus;
[0164] The data analysis component accesses the data management component based on the first identifier or the second identifier to obtain a query result fed back by the data management component, and the query result represents a duplication of the text corpus and the content corpus.
[0165] In one embodiment, the data analysis component accesses the data management component based on the first identifier or the second identifier to obtain a query result fed back by the data management component, including:
[0166] The data analysis component sends a first query instruction to the data management component based on the first identifier;
[0167] In response to the first query instruction, the data management component queries a first index based on the first identifier to feed back a first query result, and the first index is an index constructed based on hash values corresponding to the title corpus of the existing text corpus in the data storage component;
[0168] The data analysis component determines that the title corpus is duplicated in response to a case that the first query result is not empty;
[0169] The data analysis component sends a second query instruction to the data management component based on the second identifier in response to a case that the first query result is empty;
[0170] In response to the second query instruction, the data management component queries a second index based on the second identifier to feed back a second query result, and the second index is an index constructed based on hash values corresponding to the content corpus of the existing text corpus in the data storage component;
[0171] The data analysis component determines that the content corpus is duplicated in response to a case that the second query result is not empty;
[0172] The data analysis component determines that neither the title corpus nor the content corpus is duplicated in response to a case that the second query result is empty.
[0173] In one embodiment, in response to the first case, the method further includes:
[0174] The data analysis component sends a third query instruction to the data management component based on text information corresponding to the text corpus to obtain a third query result fed back by the data management component, and the third query result represents whether similar corpus corresponding to the text corpus exists in the data storage component;
[0175] In response to the third query result indicating that the similar corpus does not exist, the data storage component stores the text corpus, and the data management component stores the text information corresponding to the text corpus, the first identifier, and the second identifier.
[0176] In response to the third query result indicating that the similar corpus exists, the data management component discards the text corpus.
[0177] In an embodiment, the data analysis component sends a third query instruction to the data management component based on the text information corresponding to the text corpus, and receives a third query result fed back by the data management component, including:
[0178] The data management component obtains a text information table, and a record in the text information table indicates the first preset number of bits of information of the first text information corresponding to the text corpus stored in the data storage component.
[0179] The data management component queries the text information table based on the first preset number of bits of information of the text information, and obtains a third query result, wherein a record in the third query result satisfies a preset requirement in terms of similarity with the first preset number of bits of information of the text information.
[0180] In an embodiment, the data management component queries the text information table based on the first preset number of bits of information of the text information, and obtains a third query result, including:
[0181] The data management component calculates a Hamming distance between the first preset number of bits of information of the text information and each record in the text information table.
[0182] The third query result is obtained according to the record satisfying the preset requirement in terms of the Hamming distance.
[0183] Further, Figure 6 A hardware structure schematic diagram of a device for implementing the method provided in the embodiments of the present application is shown, and the device can participate in constituting or containing the apparatus or system provided in the embodiments of the present application. As shown in the figure, Figure 6 The device 10 can include one or more processors 102 (the processor 102 can include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. Those skilled in the art can understand that,Figure 6 The illustrated configuration is merely an example and does not limit the structure of the electronic device described above. For example, the device 10 can include more or fewer components than those shown in FIG. 1, or have a different configuration of components than those shown in FIG. 1. Figure 6 Figure 6 It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be generally referred to herein as "data processing circuitry." The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single independent processing module or any one of the other elements incorporated into the device 10 (or mobile device) in whole or in part. As referred to in embodiments of the present application, the data processing circuitry serves as a processor to control, for example, selection of a variable resistance terminal path connected to an interface.
[0184] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the method described above in embodiments of the present application. The processor 102 can execute various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e., implement the text corpus processing method described above. The memory 104 can include a high-speed random access memory and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory disposed remotely with respect to the processor 102, which can be connected to the device 10 through a network. Examples of the network can include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0185] The transmission device 106 is configured to receive or send data via a network. Examples of the network can include a wireless network provided by a communication provider of the device 10. In one example, the transmission device 106 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.
[0186] The display can be, for example, a touch screen type liquid crystal display (LCD) that can enable a user to interact with a user interface of the device 10 (or mobile device).
[0187] The display can be, for example, a touch screen type liquid crystal display (LCD) that can enable a user to interact with a user interface of the device 10 (or mobile device).
[0188] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. And the above-mentioned embodiments of the present application are described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be executed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.
[0189] Each of the embodiments in the embodiments of the present application is described in a progressive manner, and the same and similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. Especially, for the device and server embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0190] A person of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program to instruct relevant hardware to complete. The above-mentioned program can be stored in a computer readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disk, etc.
[0191] The above-mentioned is only the preferred embodiment of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A text corpus processing method, characterized in that, The method is applied to a distributed storage system, the distributed storage system including a data analysis component, and the method includes: Obtain text corpus; The data analysis component extracts the title corpus and content corpus from the text corpus; In response to the first condition, the data analysis component performs word segmentation on the text corpus to obtain the word sequence corresponding to the text corpus; the first condition indicates that both the title corpus and the content corpus are not duplicates of the corpus already stored in the distributed storage system, and the text corpus is the corpus that conforms to the first condition, determined by the data analysis component based on the first identifier and the second identifier in sequence; wherein, the first identifier is the hash value corresponding to the title corpus, and the second identifier is the hash value corresponding to the content corpus; In response to other situations different from the first situation, the data analysis component discards the text corpus; The word sequence is processed by information extraction to obtain feature information and weight information corresponding to each word in the word sequence. The weight information is determined according to the semantic importance and positional importance of the word in the word sequence. The feature vector corresponding to each word is hashed and scattered to obtain the first encoded string corresponding to each word. The first encoded string is a data string composed of binary numbers. The 0 value in the first encoded string is set to a preset negative value to obtain the encoding information corresponding to each word; Based on the encoding information and weight information corresponding to each word, the weighted encoding information corresponding to each word is obtained; The encoded information corresponding to each word is fused to obtain the text information corresponding to the text corpus.
2. The method according to claim 1, characterized in that, The process of fusing the encoded information corresponding to each word to obtain the text information corresponding to the text corpus includes: The fused encoding information is obtained by performing a bitwise addition operation on the encoding information corresponding to each word; The dimensionality reduction operation is performed on the fused encoded information to obtain the text information corresponding to the text corpus.
3. The method according to claim 1, characterized in that, The distributed storage system further includes a data storage component and a data management component. Before the operation of segmenting the text corpus to obtain the word sequence corresponding to the text corpus, the method further includes: The data analysis component acquires the first identifier or the second identifier; The data analysis component accesses the data management component based on the first identifier or the second identifier to obtain query results fed back by the data management component. The query results represent the duplication status of the text corpus and the content corpus.
4. The method according to claim 3, characterized in that, The data analysis component accesses the data management component based on the first identifier or the second identifier to obtain the query results returned by the data management component, including: The data analysis component sends a first query instruction to the data management component based on the first identifier; In response to the first query instruction, the data management component queries the first index based on the first identifier and returns the first query result. The first index is an index constructed based on the hash values corresponding to the title data of the existing text data in the data storage component. The data analysis component determines that the title corpus is duplicated when the first query result is not empty. In response to the first query result being empty, the data analysis component issues a second query instruction to the data management component based on the second identifier; In response to the second query instruction, the data management component queries the second index based on the second identifier and returns the second query result. The second index is an index constructed based on the hash values corresponding to the content corpus of the existing text corpus in the data storage component. The data analysis component determines that the content corpus is duplicated when the second query result is not empty. When the second query result is empty, the data analysis component determines that neither the title corpus nor the content corpus is repeated.
5. The method according to claim 3, characterized in that, In response to the first case, the method further includes: The data analysis component sends a third query instruction to the data management component based on the text information corresponding to the text corpus, so as to obtain the third query result fed back by the data management component. The third query result indicates whether there is similar corpus corresponding to the text corpus in the data storage component. In response to the third query result indicating that no similar corpus exists, the text corpus is stored in the data storage component, and the text information corresponding to the text corpus, the first identifier, and the second identifier are stored in the data management component; In response to the third query result indicating the existence of similar corpus, the text corpus is discarded.
6. The method according to claim 5, characterized in that, Based on the text information corresponding to the text corpus, the data analysis component sends a third query instruction to the data management component to obtain a third query result returned by the data management component, including: The data management component obtains a text information table, and the records in the text information table represent the first preset number of bits of information corresponding to the first text information that has been stored in the data storage component. The data management component queries the text information table based on the information of the first preset number of positions of the text information to obtain a third query result. The similarity between the records in the third query result and the information of the first preset number of positions of the text information meets the preset requirements.
7. The method according to claim 6, characterized in that, The data management component queries the text information table based on the information of the first preset number of digits of the text information to obtain a third query result, including: The data management component calculates the information of the first preset number of bits of the text information and the Hamming distance between them and each record in the text information table; The third query result is obtained based on the records whose Hamming distance meets the preset requirements.
8. A text corpus processing device, characterized in that, The device is deployed in a distributed storage system, which includes a data analysis component, and the device includes: The text corpus acquisition module is used to acquire text corpora. The data analysis component is used to extract title data and content data from the text corpus; In response to the first condition, the data analysis component is used to perform word segmentation on the text corpus to obtain the word sequence corresponding to the text corpus; the first condition indicates that both the title corpus and the content corpus are not duplicates of the corpus already stored in the distributed storage system, and the text corpus is the corpus that conforms to the first condition, determined by the data analysis component based on the first identifier and the second identifier in sequence; wherein, the first identifier is the hash value corresponding to the title corpus, and the second identifier is the hash value corresponding to the content corpus; In response to other situations different from the first situation, the data analysis component is used to discard the text corpus; The information extraction module is used to perform information extraction processing on the word sequence to obtain feature information and weight information corresponding to each word in the word sequence. The weight information is determined according to the semantic importance and positional importance of the word in the word sequence. The hash module is used to perform hash scattering on the feature vector corresponding to each word to obtain a first encoding string corresponding to each word, wherein the first encoding string is a data string composed of binary numbers; and is used to set the 0 value in the first encoding string to a preset negative value to obtain the encoding information corresponding to each word. The weighting module is used to obtain the weighted encoding information corresponding to each word based on the encoding information and weight information corresponding to each word; The fusion module is used to fuse the encoded information corresponding to each word to obtain the text information corresponding to the text corpus.
9. A distributed storage system, characterized in that, The system includes a text corpus processing apparatus as described in claim 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement a text corpus processing method as described in any one of claims 1 to 7.
11. An electronic device, characterized in that, The method includes at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements a text corpus processing method as described in any one of claims 1 to 7 by executing the instructions stored in the memory.
12. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they implement a text corpus processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Duplicate detection system and method for search engines
CN102375813A
Microblog duplication-eliminating method and system based on reverse-order index
CN103646080A
Media text similarity detection method
CN113111645A