Information processing method, device and computer readable storage medium

Through the methods of word segmentation processing and sentence vector analysis, the problem of low text deduplication efficiency and accuracy in the prior art is solved, automatic and accurate text deduplication is achieved, and the efficiency and accuracy of information processing are improved.

CN114328885BActive Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111485271.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-05-23
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

When the prior art faces massive text to be removed, manual methods are time-consuming and text algorithms can only judge whether the composition is repeated and cannot semantically deduplicate, resulting in low information processing efficiency and accuracy.

Method used

By obtaining the target pending text set, word segmentation process generates word segmentation sets of different word lengths, calculates sentence vectors, removes principal component vectors, calculates the similarity between sentence vectors, and deduplication processing according to preset thresholds.

Benefits of technology

Automatic and accurate text deduplication is achieved, the efficiency and accuracy of information processing are improved, and the semantics of different statements can be better distinguished.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114328885B_ABST
    Figure CN114328885B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses an information processing method, device and computer-readable storage medium. The embodiment of the present application obtains a target set of texts to be processed; each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed; the sentence vector corresponding to each target text to be processed is calculated based on the word vector corresponding to each segmentation set; the principal component vector in each sentence vector is removed to obtain multiple target sentence vectors after the removal; the similarity between each target sentence vector is calculated, and the target text to be processed whose similarity is less than a preset threshold is deduplicated. In this way, by removing the principal component vector, the deduplication of big data is achieved efficiently and accurately. The technical solution of the embodiment of the present application can be applied to cloud computing, maps, big data, artificial intelligence and other fields, improving the efficiency and accuracy of information processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information processing technology, and in particular to an information processing method, device and computer-readable storage medium. Background Art

[0002] With the development of the Internet and the widespread application of computers, the Internet is filled with a large amount of repeated text content, especially in some business areas such as training data and advertising. If there is a large amount of repeated text content, it will not only reduce the overall text quality, but also waste a lot of storage resources.

[0003] In the prior art, in order to save storage resources, it is necessary to remove duplicate text content, for example, by manually comparing multiple texts in pairs to remove duplicate texts, or by comparing the similarities between texts through some text algorithms to remove similar texts to achieve the effect of deduplication.

[0004] In the process of research and practice of the prior art, the inventors of the present application found that in the prior art, when faced with massive amounts of text to be removed, the manual method would waste a lot of time, and the text algorithm comparison method could often only determine whether the structure of the text was repeated, but could not deduplicate semantically, resulting in low efficiency and accuracy of information processing. Summary of the invention

[0005] The embodiments of the present application provide an information processing method, device, and computer-readable storage medium, which can improve the efficiency and accuracy of information processing.

[0006] To solve the above technical problems, the present application provides the following technical solutions:

[0007] An information processing method, comprising:

[0008] Acquire a target text set to be processed, wherein the target text set to be processed includes a plurality of target texts to be processed;

[0009] Each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed;

[0010] Calculate the sentence vector of each target text to be processed based on the word vector corresponding to each word segmentation set;

[0011] The principal component vector in each sentence vector is removed to obtain multiple target sentence vectors after the removal process;

[0012] The similarity between each target sentence vector is calculated, and the target sentence vectors with similarity less than a preset threshold are deduplicated to the corresponding target text pairs to be processed.

[0013] An information processing device, comprising:

[0014] An acquisition unit, used for acquiring a target text set to be processed, wherein the target text set to be processed includes a plurality of target texts to be processed;

[0015] The word segmentation unit is used to perform word segmentation processing on each target text to be processed according to different word lengths in turn, and obtain a word segmentation set corresponding to each word length of each target text to be processed;

[0016] A first calculation unit is used to calculate a sentence vector corresponding to each target text to be processed based on the word vector corresponding to each word segmentation set;

[0017] A removal unit, used for removing the principal component vector in each sentence vector to obtain a plurality of target sentence vectors after the removal process;

[0018] The second calculation unit is used to calculate the similarity between each target sentence vector, and perform deduplication processing on the target sentence vector with a similarity less than a preset threshold and the corresponding target text to be processed.

[0019] In some embodiments, the removal unit comprises:

[0020] A combination subunit, used to combine each sentence vector to obtain a sentence vector matrix;

[0021] An analysis subunit, used for performing principal component analysis on the sentence vector matrix to obtain a principal component vector matrix;

[0022] The removal subunit is used to remove each sentence vector in the sentence vector matrix from the principal component vector matrix in turn to obtain a target sentence vector matrix, wherein the target sentence vector matrix includes multiple target sentence vectors.

[0023] In some embodiments, the removal subunit is used to:

[0024] Obtaining a transposed matrix corresponding to the principal component vector matrix;

[0025] The difference between each sentence vector in the sentence vector matrix and the product of the principal component vector matrix, the transposed matrix and the corresponding sentence vector is calculated to obtain a calculated target sentence vector matrix.

[0026] In some embodiments, the second computing unit is configured to:

[0027] Split the calculated target sentence vector matrix to obtain multiple target sentence vectors;

[0028] The cosine similarity between each target sentence vector is calculated, and the target sentence vectors whose cosine similarity is less than a preset cosine threshold are deduplicated to the corresponding target text pairs to be processed.

[0029] In some embodiments, the acquisition unit is used to:

[0030] Obtaining a text set to be processed, wherein the text set to be processed includes a plurality of texts to be processed;

[0031] The stop words in each to-be-processed text are removed to obtain a plurality of target to-be-processed texts after the removal to generate a target to-be-processed text set.

[0032] In some embodiments, the word segmentation unit is used to:

[0033] Each target text to be processed is segmented in turn according to the sliding windows corresponding to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed.

[0034] A computer-readable storage medium stores a plurality of instructions, wherein the instructions are suitable for a processor to load to execute the steps in the above-mentioned information processing method.

[0035] A computer program product or a computer program, wherein the computer program product or the computer program comprises computer instructions, wherein the computer instructions are stored in a storage medium. A processor of a computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions, so that the computer performs the steps in the above-mentioned information processing method.

[0036] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the above-mentioned information processing method when executing the computer program.

[0037] The embodiment of the present application obtains a target set of texts to be processed; each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed; the sentence vector corresponding to each target text to be processed is calculated based on the word vector corresponding to each segmentation set; the principal component vector in each sentence vector is removed to obtain multiple target sentence vectors after the removal; the similarity between each target sentence vector is calculated, and the target sentence vectors with similarity less than a preset threshold are deduplicated to the corresponding target text to be processed. In this way, a sentence vector with accurate semantic expression is generated through word segmentation, and the principal component vector of the sentence vector is removed to make the difference between the vectors more obvious, so that different sentences can be better distinguished when judging the similarity of sentences. Compared with the existing manual text deduplication method, the embodiment of the present application can realize an automatic and accurate text deduplication method, which improves the efficiency and accuracy of information processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0039] Figure 1 is a schematic diagram of a scenario of an information processing system provided by an embodiment of the present application;

[0040] Figure 2 It is a flowchart of the information processing method provided in the embodiment of the present application;

[0041] Figure 3 is another flowchart of the information processing method provided by an embodiment of the present application;

[0042] Figure 4 A schematic diagram of the structure of an open source cluster computing framework provided in an embodiment of the present application;

[0043] Figure 5 Another flowchart of the information processing method provided by the embodiment of the present application is

[0044] Figure 6 is a schematic diagram of the structure of an information processing device provided in an embodiment of the present application;

[0045] Figure 7 It is a schematic diagram of the structure of the server provided in the embodiment of the present application. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0047] Embodiments of the present application provide an information processing method, device, and computer-readable storage medium.

[0048] See also Figure 1 , Figure 1 The scenario diagram of the information processing system provided for the embodiment of the present application includes: a terminal, and a server (the information processing system may also include other terminals in addition to the terminal, and the specific number of terminals is not limited here), the terminal and the server may be connected through a communication network, and the communication network may include a wireless network and a wired network, wherein the wireless network includes a wireless wide area network, a wireless local area network, a wireless metropolitan area network, and a combination of one or more of a wireless personal network. The network includes network entities such as routers and gateways, which are not shown in the figure. The terminal can exchange information with the server through the communication network. For example, when the terminal is running an application containing various types of push information, such as video, short video, microblog, and advertising applications, the terminal can send the target text set to be processed that needs to be deduplicated to the server.

[0049] The information processing system may include an information processing device, which may be integrated in a server. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Figure 1 In the method, the server is mainly used to obtain a target text set to be processed, which contains multiple target texts to be processed; each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed; a sentence vector corresponding to each target text to be processed is calculated based on the word vector corresponding to each segmentation set; the principal component vector in each sentence vector is removed to obtain multiple target sentence vectors after the removal; the similarity between each target sentence vector is calculated, and the target sentence vectors with a similarity less than a preset threshold are deduplicated to the corresponding target texts to be processed, and the deduplicated target text set is sent to the terminal for display.

[0050] The information processing system may also include a terminal, which can install various applications required by users, such as videos, short videos, microblogs, advertisements and other applications. For example, the terminal can send the target text set to be processed that needs to be deduplicated to the server, and can also receive the target text set to be processed after deduplication returned by the server. Since duplicate text is removed, storage resources can be saved, and the overall text quality can be improved, so that better subsequent processing effects can be achieved.

[0051] It should be noted that Figure 1 The scenario diagram of the information processing system shown is merely an example. The information processing system and scenario described in the embodiments of the present application are intended to more clearly illustrate the technical solution of the embodiments of the present application, and do not constitute a limitation on the technical solution provided in the embodiments of the present application. A person of ordinary skill in the art will appreciate that with the evolution of information processing systems and the emergence of new business scenarios, the technical solution provided in the embodiments of the present application is equally applicable to similar technical problems.

[0052] The following are detailed descriptions of each.

[0053] In this embodiment, the description will be made from the perspective of an information processing device, which can be specifically integrated into a server having a storage unit and a microprocessor installed therein and having computing capabilities.

[0054] See also Figure 2 , Figure 2 : is a flow chart of an information processing method provided in an embodiment of the present application. The information processing method includes:

[0055] In step 101, a target text set to be processed is obtained.

[0056] Among them, the target to-be-processed text set contains multiple target to-be-processed texts, and each target to-be-processed text can be understood as a sentence. In the relevant technology, in the actual recommendation business, such as text recommendations such as articles, advertisements, or news, in order to achieve better recommendation effects, some repeated texts need to be deleted. Moreover, with the development of artificial intelligence, some repeated training data, such as repeated user portraits, will lead to slower training efficiency. It is understandable that in the specific implementation of the present application, related data such as user portraits are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. And the stacking of repeated content will also lead to a waste of server storage space and increase unnecessary costs.

[0057] In order to solve the above problem, the embodiment of the present application can obtain a target text set to be processed, which can contain multiple texts to be processed, such as 1,000 or 10,000 texts, and each target text to be processed is a sentence, such as "Tom chases Jerry".

[0058] In one embodiment, the embodiment of the present application can obtain the target text set to be processed through cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model application, which can form a resource pool, which is used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, each item may have its own identification mark in the future, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately, and all kinds of industry data require strong system backing support, which can only be achieved through cloud computing.

[0059] In some implementations, obtaining a target text set to be processed may include:

[0060] (1) obtaining a text set to be processed, wherein the text set to be processed includes a plurality of texts to be processed;

[0061] (2) The stop words in each to-be-processed text are removed to obtain a plurality of target to-be-processed texts after the stop words are removed to generate a target to-be-processed text set.

[0062] Among them, a set of texts to be processed can be obtained, which contains multiple texts to be processed. The texts to be processed are unoptimized texts, such as "Tom likes to chase Jerry~". The number can be 1,000 or 10,000, and there is no limit on the number here.

[0063] Furthermore, the stop words may include English characters, numbers, mathematical characters, punctuation marks and single Chinese characters with extremely high usage frequency. Since the stop words are of no substantial help to the expression of sentences, the embodiment of the present application can optimize the removal of stop words in each text to be processed, and obtain multiple target texts to be processed after removal to generate a target text set to be processed. For example, "Tom likes to chase Jerry~" is optimized to "Tom likes to chase Jerry". In this way, the multiple target texts to be processed after removal can express accurate meanings with fewer words, reduce subsequent calculations, and improve the efficiency of information processing.

[0064] In step 102, each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed.

[0065] In the related art, directly judging the similarity between the target texts to be processed through a hash algorithm can often only judge whether the structure of the texts is repeated, but cannot judge whether the semantics of the texts are repeated. In the actual processing process, texts with different structures can actually express the same meaning. For example, "Tom chases Jerry" and "Jerry is chased by Tom" are different in structure, but the semantics are the same. The related art cannot achieve deduplication in this scenario, which will result in poor deduplication effect.

[0066] In an embodiment of the present application, an algorithm of a statistical language model (N-Gram) can be used to implement word segmentation of each text to be processed according to different word lengths. The specific implementation method includes: performing a sliding window operation of size N on the content of the sentence according to bytes to form a byte segment sequence of length N. The model is based on the Markov hypothesis, that is, the appearance of the Nth word is only related to the previous N-1 words, and is not related to any other words. The probability of the entire sentence is the product of the probabilities of occurrence of each word. Based on this idea, the target sentence vector of each target text to be processed can be calculated later.

[0067] In this way, each target text to be processed can be segmented according to different word lengths, such as 2 words, 3 words, etc., to obtain a segmentation set corresponding to each word length of each target text to be processed. For example, target text to be processed 1 can have a segmentation set corresponding to 2 words and 3 words, and target text to be processed 2 can also have a segmentation set corresponding to 2 words and 3 words. And so on, each target text to be processed has a corresponding segmentation set for each word length. The number of word lengths can be set by the user, such as word lengths of 2, 3, and 4 or word lengths of 2, 3, 4, and 5, and can also be configured according to different application scenarios, which is not specifically limited here.

[0068] In some implementations, each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed, including:

[0069] (1) Each target text to be processed is segmented in turn according to the sliding windows corresponding to different word lengths, and a segmentation set corresponding to each word length of each target text to be processed is obtained.

[0070] Among them, the N-Gram can include binary 2-Gram and ternary 3-Gram, etc. The 2-Gram represents a sliding window corresponding to a word length of 2, and the 3-Gram represents a sliding window corresponding to a word length of 3. Correspondingly, the N-Gram represents a sliding window corresponding to a word length of N. Based on this, each target text to be processed Si (i represents the number of texts) can be segmented in turn according to the sliding windows corresponding to different word lengths, and the 2-Gram words of the same target text to be processed are placed in the word set b2, the 3-Gram words are placed in the word set b3, and the N-Gram words are placed in the word set bn. Therefore, each target text to be processed is split into an N-Gram data set S i ={b2, b3,…, bn}.

[0071] For example, the target text to be processed is "Tom chases Jerry". According to the sliding window corresponding to 2-Gram, word segmentation processing can be performed, and the word segmentation set corresponding to the word length of 2 for the target text to be processed "Tom chases Jerry" can be obtained (Tom, Tom chases, Jerry chases, Jerry). And so on, the word segmentation set corresponding to each word length of the target text to be processed "Tom chases Jerry" can be obtained.

[0072] In step 103, a sentence vector corresponding to each target text to be processed is calculated based on the word vector corresponding to each word segmentation set.

[0073] In order to obtain the corresponding sentence vector of each target text to be processed, the word segments in the same target text to be processed can be first converted into corresponding vectors. The vector can be understood as a vector of word to word conversion, that is, mapping a word to a new space and representing it with a multi-dimensional continuous real number vector is called word embedding, which is a general term for a group of language modeling and feature learning technologies in natural language processing (NLP). In which words or phrases from the vocabulary are mapped to real number vectors. It involves mathematical embedding from a one-dimensional space for each word to a continuous vector space with a lower dimension.

[0074] In one embodiment, the embodiment of the present application can convert each word in each word segmentation set into a corresponding vector through a language model algorithm, such as the word2vec algorithm, for example, converting the word "Tom" into [0.01, 0.23, 0.89, ... 0.92], and generating a word vector dictionary corresponding to all word segments. The word2vec algorithm can represent words as an efficient algorithm model for real-valued vectors. It uses the idea of ​​deep learning and can simplify the processing of text content into vector operations in a K-dimensional vector space through training. The similarity in the vector space can be used to represent the semantic similarity of the text.

[0075] Furthermore, the word2vec algorithm can be used to convert the words in each word segmentation set in the same target text to be processed into vectors, and then the vectors corresponding to all the words in the same word segmentation set are combined to obtain the word vector corresponding to each word segmentation set. The word vector can express the word segmentation sentence vector expressed by the word segmentation set with corresponding word lengths. Since the word segmentation sentence vector is composed of vector expressions corresponding to different word segmentations, different sentences with similar semantics can be identified. Moreover, since the same target text to be processed contains multiple word segmentation sets, the embodiment of the present application can continue to combine the word segmentation sentence vectors corresponding to multiple word segmentation sets of the same target text to be processed to form a sentence vector corresponding to the target text to be processed. Since the sentence vector further integrates the word segmentation sentence vectors of word segmentation sets of multiple word lengths, it will express the semantics of the target text to be processed more accurately, and subsequent deduplication can be more accurate.

[0076] In one embodiment, the calculation of the sentence vector corresponding to each target text to be processed based on the word vector corresponding to each word segmentation set may include:

[0077] (1) Calculate the word vector of each word segmentation set corresponding to each target text to be processed in turn;

[0078] (2) According to the word length corresponding to each word segmentation set, different weights are set for each word vector;

[0079] (3) Calculate each word vector and the corresponding weight of the same target text to be processed to obtain the corresponding sentence vector of each target text to be processed.

[0080] Among them, the vector corresponding to each word in each word segmentation set corresponding to each target to-be-processed text can be calculated in turn by the word2vec algorithm, and then the vector corresponding to each word in the same word segmentation set in the same target to-be-processed text can be counted to obtain the word vector of each word segmentation set in the same target to-be-processed text. It is easy to understand that in the scenario of word segmentation with different word lengths, the longer the word length, the lower the word frequency, and the shorter the word length, the higher the word frequency. In practical applications, the importance of words with greater word frequency is often lower, and on the contrary, the importance of words with smaller word frequency is often higher. In this way, according to the word length corresponding to each word segmentation set, the corresponding weight can be set for the word vector corresponding to each word segmentation set. The smaller the word length corresponding to the word segmentation set, the smaller the corresponding weight is set. On the contrary, the larger the word length corresponding to the word segmentation set, the larger the corresponding weight is set. For example, a weight of 0.2 can be set for a word segmentation set of 2 words, a weight of 0.3 can be set for a word segmentation set of 3 words, and a weight of 0.5 can be set for a word segmentation set of 4 words.

[0081] Furthermore, each word vector of the same target text to be processed can be adjusted according to the corresponding weight, and each word vector after adjustment can be combined to obtain the corresponding sentence vector of each target text to be processed.

[0082] In step 104, the principal component vector in each sentence vector is removed to obtain a plurality of target sentence vectors after the removal process.

[0083] Among them, there may be correlations between different target texts to be processed. For example, there is common information of "is" between the target text to be processed "Tom is a bad man" and the target text to be processed "Tom is a good man". As a result, there is common information between the sentence vectors after the target text to be processed is transformed. In order to achieve better subsequent judgment of the similarity between the two target texts to be processed, the common information can be removed. The common information can be a principal component vector. In one embodiment, the principal component vector can be obtained by performing principal component analysis (PCA) on all sentence vectors. The principal component analysis method is one of the most widely used data dimensionality reduction algorithms. The main idea of ​​PCA is to map n-dimensional features to k-dimensions, which are smaller than n-dimensions. This k-dimension is a new orthogonal feature also called a principal component. It is a k-dimensional feature reconstructed on the basis of the original n-dimensional feature, which can be used to extract the main feature components of the data.

[0084] The principal component vector is an important principal component vector analyzed from all the sentence vectors, that is, the common information corresponding to all the sentence vectors. The embodiment of the present application can remove the principal component vector from each sentence vector, that is, remove the common information in each sentence vector to obtain multiple target sentence vectors after the removal process. Since the common information is removed from the target sentence vectors, the difference expression between the target sentence vectors is more obvious, which can make the subsequent deduplication processing more accurate.

[0085] In one embodiment, removing the principal component vector in each sentence vector to obtain multiple target sentence vectors after the removal process may include:

[0086] (1) Combine each sentence vector to obtain a sentence vector matrix;

[0087] (2) performing principal component analysis on the sentence vector matrix to obtain a principal component vector matrix;

[0088] (3) The principal component vector matrix is ​​removed from each sentence vector in the sentence vector matrix in turn to obtain a target sentence vector matrix, which contains multiple target sentence vectors.

[0089] Among them, each sentence vector W can be obtainedi , the entire sentence vector W i The Reduce operator is used to perform specific calculations on a certain dimension of multidimensional tensor data to achieve the purpose of reducing the dimension. In this way, the sentence vector matrix X can be obtained, X = {W 1 , W 2 , W 3 , …, W i}, the dimension of the matrix is ​​(i, d), where n is the number of all texts and d is the size of each sentence vector.

[0090] Furthermore, the matrix X is subjected to principal component analysis by the PCA tool to obtain the principal component matrix μ, which is an important component matrix of the matrix X after dimensionality reduction, and can also be understood as the common information of the matrix X. Thus, each sentence vector in the sentence vector matrix can be removed from the principal component vector matrix in turn to obtain a target sentence vector matrix with the common information removed, and the target sentence vector matrix contains multiple target sentence vectors.

[0091] In one embodiment, removing the principal component vector matrix from each sentence vector in the sentence vector matrix in turn to obtain a target sentence vector matrix may include:

[0092] (1.1) Obtain the transposed matrix corresponding to the principal component vector matrix;

[0093] (1.2) Calculate the difference between each sentence vector in the sentence vector matrix and the product of the principal component vector matrix, the transposed matrix and the corresponding sentence vector to obtain the calculated target sentence vector matrix.

[0094] Among them, the transposed matrix corresponding to the principal component vector matrix can be obtained. The transposed matrix is ​​a new matrix obtained by swapping the rows and columns of the matrix, and the determinant of the transposed matrix remains unchanged. In order to better describe the calculation process, please refer to the following formula:

[0095]

[0096] Among them, the is the target sentence vector matrix, μ is the principal component matrix, μ T is the transposed matrix of the principal component matrix. Based on this, the above formula is used to calculate each sentence vector w in the sentence vector matrix. i With the principal component matrix μ, transposed matrix μ T And the corresponding sentence corresponding to w i The difference of the product of , and the calculated target sentence vector matrix is ​​obtained.

[0097] In step 105, the similarity between each target sentence vector is calculated, and the target sentence vectors with similarity less than a preset threshold are deduplicated for the corresponding target text pairs to be processed.

[0098] Among them, the closer the spatial distance between different target sentence vectors, the more similar they are, and the farther the spatial distance, the less similar they are. In one embodiment, the similarity can be calculated by Euclidean distance or cosine similarity, and the preset threshold is the critical value that determines whether the two target texts to be processed corresponding to each two target sentence vectors are the same text.

[0099] Therefore, since common information is removed from each target sentence vector, the expression differences between the target sentence vectors are more obvious, and the calculated similarity between each target sentence vector is more accurate. The target text pairs corresponding to the target sentence vector pairs whose similarity is less than a preset threshold can be judged as the same text, and any target text to be processed in the text pairs to be processed that are judged to be the same can be deleted to achieve accurate deduplication. Since the semantics of the expression of the target sentence vector in the embodiment of the present application is more accurate, and the differences between the target sentence vectors are made more significant by removing the principal component vectors, the similarity calculation can be made more accurate, and a better deduplication effect can be achieved.

[0100] As can be seen from the above, the embodiment of the present application obtains a target text set to be processed; each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed; the sentence vector corresponding to each target text to be processed is calculated based on the word vector corresponding to each segmentation set; the principal component vector in each sentence vector is removed to obtain multiple target sentence vectors after the removal; the similarity between each target sentence vector is calculated, and the target sentence vector with a similarity less than a preset threshold is deduplicated for the corresponding target text to be processed. In this way, a sentence vector with accurate semantic expression is generated through word segmentation, and the principal component vector of the sentence vector is removed to make the difference between the vectors more obvious, so that different sentences can be better distinguished when judging the similarity of sentences. Compared with the existing manual text deduplication method, the embodiment of the present application can realize an automatic and accurate text deduplication method, which improves the efficiency and accuracy of information processing.

[0101] The following is a further detailed description with examples.

[0102] In this embodiment, the information processing device is specifically integrated into a server as an example for description.

[0103] See also Figure 3 , Figure 3 Another flowchart of the information processing method provided in the embodiment of the present application. The method flow may include:

[0104] In step 201, the server obtains a set of texts to be processed, removes stop words in each text to be processed, and obtains a plurality of target texts to be processed after the removal to generate a target set of texts to be processed.

[0105] The server in the embodiment of the present application may be a cloud server using cloud technology, and may integrate an open source cluster computing framework, such as the Spark engine, which is a fast general computing engine designed for big data processing. Figure 4 As shown, Figure 4 A schematic diagram of the structure of an open source cluster computing framework provided in an embodiment of the present application.

[0106] The application layer A may include a package for structured data (Spark SQL), a component for stream computing (Spark Streaming), a library for machine learning (MLlib (machine learning)), and a tool set for graph operations and calculations (Graph X). The package for structured data is a package used by Spark to operate structured data. Through Spark SQL, SQL dialects can be used to query data. Spark SQL supports multiple data sources, such as data warehouse tools (Hive) tables, etc. The component for stream computing is a component provided by Spark for streaming real-time data, and an application programming interface (Application Programming Interface, API) for operating data streams is provided. The library for machine learning provides a library of common machine learning functions, including classification, regression, clustering, collaborative filtering, etc., and also provides additional support functions such as model evaluation and data import. The tool set for graph operations and calculations is a set of algorithms and tools for controlling graphs, parallel graph operations and calculations.

[0107] The core data computing layer B may include the code function layer (Spark Core) of the open source cluster computing framework, which implements the basic functions of Spark, including modules such as task scheduling, memory management, error recovery and storage system interaction. The Spark Core also includes the API definition of Resilient Distributed Datasets (RDD).

[0108] The resource scheduling layer C may include a local operation mode, an open source general resource management system (YARN), an open source distributed resource management framework (Mesos), etc., for resource management.

[0109] The data resource layer D may include a distributed file system (Hadoop Distributed File System, HDFS) or a distributed, column-oriented open source database (HBase), etc.

[0110] The Spark engine mentioned above can realize distributed iterative processing of data, provide computing speed for efficient processing of data streams, and the Spark supports APIs in multiple development languages, which can quickly build different applications.

[0111] In this way, the server can load the set of texts to be processed on the distributed file system through the Spark engine and implement subsequent calculation processing. The set of texts to be processed contains multiple texts to be processed. For example, a text to be processed is "Little Tom likes to chase Jerry~", and the number can be 1,000 or 10,000. There is no limit on the number here.

[0112] Furthermore, the stop words in each text to be processed can be removed for optimization, and the multiple target texts to be processed after the removal are obtained to generate a target text set to be processed. For example, "Little Tom likes to chase Jerry~" is optimized to "Tom likes to chase Jerry". In this way, the multiple target texts to be processed after the removal can express the accurate meaning with fewer words, reduce the subsequent calculation amount, and improve the efficiency of information processing.

[0113] In step 202, the server performs word segmentation processing on each target text to be processed in turn according to the sliding windows corresponding to different word lengths, and obtains a word segmentation set corresponding to each word length of each target text to be processed.

[0114] Among them, it is assumed that the target text to be processed is S i , i is the target text to be processed, and each target text to be processed can be segmented according to the sliding windows corresponding to N different word lengths corresponding to N-Gram. For example, N=4, 2-Gram represents the sliding window corresponding to the word length of 2, 3-Gram represents the sliding window corresponding to the word length of 3, and 4-Gram represents the sliding window corresponding to the word length of 4. Based on this, the 2-Gram words of the same target text to be processed can be put into the word set b2, the 3-Gram words into the word set b3, and the 4-Gram words into the word set b4. Therefore, each target text to be processed is split into an N-Gram data set S i ={b2, b3, ..., bn}, n = 4. That is, each target text to be processed can correspond to 4 word segmentation sets.

[0115] In step 203, the server obtains the vector and word frequency information corresponding to each word in each word segmentation set, calculates the vector and corresponding word frequency information of each word in each word segmentation set in the same target text to be processed, and obtains the word vector of each word segmentation set corresponding to each target text to be processed.

[0116] In practical applications, the importance of words with higher frequency is often lower, while the importance of words with lower frequency is often higher. Therefore, it is necessary to obtain the vector and frequency information corresponding to each word in each word set. The vector information can be calculated for each word through the word2vec algorithm. The frequency information is the probability of the word appearing in all words. For example, if the word "Jerry" appears 20 times in all 100 words, the frequency information of the word "Jerry" is 0.2.

[0117] Furthermore, in order to better illustrate the embodiments of the present application, the following formula may be referred to:

[0118]

[0119] Among them, the g 2 is the word vector of the word segmentation set corresponding to any target text to be processed when the word length is 2. α is a hyperparameter, that is, a known parameter. w is the vector of each word segment in the hierarchical set with a word length of 2, w is the number of word segments, and the p w is the frequency of each word segment. Based on the above formula, the vector v of each word segment is w Divide by the frequency of the segmented word and sum them up, so that the segmented words with higher frequency have lower proportion, and the segmented words with lower frequency have higher proportion, which is in line with practical applications. Finally, we get the word vector that better expresses the meaning of each segmented word set. By analogy, we can calculate the word vector of the segmented word set corresponding to the word length of 3 and the word vector of the segmented word set corresponding to the word length of 4.

[0120] In step 204, the server sets different weight information for each word vector according to the word length corresponding to each word segmentation set.

[0121] Among them, the corresponding weight can be set for the word vector corresponding to each word segmentation set according to the word length corresponding to each word segmentation set. The smaller the word length corresponding to the word segmentation set, the smaller the corresponding weight is set. On the contrary, the larger the word length corresponding to the word segmentation set, the larger the corresponding weight is set. For example, a weight of 0.2 can be set for a word segmentation set of 2 words, a weight of 0.3 can be set for a word segmentation set of 3 words, and a weight of 0.5 can be set for a word segmentation set of 4 words.

[0122] In step 205, the server multiplies each word vector of each target text to be processed by the corresponding weight to obtain multiple products corresponding to each target text to be processed, and sums the multiple products corresponding to the same target text to be processed to obtain the sentence vector corresponding to each target text to be processed.

[0123] In order to better illustrate the embodiments of the present application, the following formula may be referred to:

[0124]

[0125] The W i is a sentence vector, the number of i is equal to the number of target texts to be processed, g i is the word vector of the corresponding word segmentation set of any target text to be processed with word length i. i is the weight of the word segmentation set corresponding to the word length i, and each word vector g of the same target text to be processed is i and the corresponding weight β i Multiply them and sum the products of the same text to be processed to obtain the sentence vector W corresponding to each target text to be processed i .

[0126] In step 206, the server combines each sentence vector to obtain a sentence vector matrix, and performs principal component analysis on the sentence vector matrix to obtain a principal component vector matrix.

[0127] Among them, each sentence vector W can be obtained i , the entire sentence vector W i Transfer to the Reduce operator for aggregation and dimension reduction, so as to obtain the sentence vector matrix X, X = {W 1 , W 2 , W 3 , …, W i}, the dimension of the matrix is ​​(i, d), where n is the number of all texts and d is the size of each sentence vector.

[0128] Furthermore, the sentence vector matrix X is subjected to principal component analysis by using the PCA tool to obtain the principal component matrix μ, which is an important component matrix of the matrix X after dimensionality reduction and can also be understood as the common information contained in the matrix X.

[0129] In step 207, the server obtains the transposed matrix corresponding to the principal component vector matrix, calculates the difference between each sentence vector in the sentence vector matrix and the product of the principal component vector matrix, the transposed matrix and the corresponding sentence vector, and obtains the calculated target sentence vector matrix.

[0130] Among them, the transposed matrix μ corresponding to the principal component vector matrix μ can be obtainedT , the transposed matrix μ T The new matrix obtained by swapping the rows and columns of the matrix is ​​called the transposed matrix. T The determinant of remains unchanged. To better describe the calculation process, please refer to the following formula:

[0131]

[0132] Among them, the is the target sentence vector matrix, μ is the principal component matrix, μ T is the transposed matrix of the principal component matrix. Based on this, the above formula is used to calculate each sentence vector w in the sentence vector matrix. i With the principal component matrix μ, transposed matrix μ T And the corresponding sentence corresponding to w i The difference of the product of , minus the principal component matrix μ, removes the principal component vector corresponding to the common information of all sentences, and obtains the calculated target sentence vector matrix The retained target sentence vector matrix It can better represent the difference between itself and other target sentence vector matrices. After the above transformation, the matrix X is converted to and

[0133] In step 208, the server splits the calculated target sentence vector matrix to obtain multiple target sentence vectors, calculates the cosine similarity between each target sentence vector, and deduplicates the target sentence vectors whose cosine similarity is less than a preset cosine threshold and the corresponding target text pairs to be processed.

[0134] The server can Split to obtain multiple target sentence vectors Each target sentence vector The corresponding text id can be identified, and the text id is associated with the corresponding target text to be processed (i.e., sentence). The closer the cosine similarities between different target sentence vectors are, the closer the semantics of the corresponding target texts to be processed are. The less similar the cosine similarities between different target sentence vectors are, the less similar the semantics of the corresponding target texts to be processed are. The preset cosine threshold is the critical value for determining whether the two target texts to be processed corresponding to the two target sentence vectors are the same text.

[0135] In this way, any two target sentence vectors can be calculated The cosine similarity between the two target texts is calculated, and the target sentence vectors whose cosine similarity is less than the preset cosine threshold are deduplicated for the corresponding target text pairs to be processed, that is, any one of the two text IDs associated with the target sentence vector pair is directly filtered. Finally, all target texts to be processed with repeated text content are deduplicated. Since the generation of the target sentence vector is combined with semantics and the common information is removed through principal component analysis, the similarity calculation between the target sentence vectors is more accurate, and a more comprehensive deduplication effect is achieved.

[0136] As can be seen from the above, the embodiment of the present application obtains a target text set to be processed; each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed; the sentence vector corresponding to each target text to be processed is calculated based on the word vector corresponding to each segmentation set; the principal component vector in each sentence vector is removed to obtain multiple target sentence vectors after the removal; the similarity between each target sentence vector is calculated, and the target sentence vector with a similarity less than a preset threshold is deduplicated for the corresponding target text to be processed. In this way, a sentence vector with accurate semantic expression is generated through word segmentation, and the principal component vector of the sentence vector is removed to make the difference between the vectors more obvious, so that different sentences can be better distinguished when judging the similarity of sentences. Compared with the existing manual text deduplication method, the embodiment of the present application can realize an automatic and accurate text deduplication method, which improves the efficiency and accuracy of information processing.

[0137] Furthermore, since different weights are set for different word segmentations and word segmentation sets according to actual applications, the expression of the target sentence vector is more consistent with the semantic expression of the sentence, further improving the accuracy of information processing.

[0138] The following will provide further detailed explanation with examples.

[0139] In this embodiment, the information processing device is specifically integrated into a server as an example for description.

[0140] See also Figure 5 , Figure 5 Another flowchart of the information processing method provided in the embodiment of the present application. The method flow may include:

[0141] In step 11, the Spark engine reads and parses the distributed file system data.

[0142] Among them, the log text data on the distributed file system can be loaded through the Spark engine to obtain a set of texts to be processed, which contains multiple texts to be processed. Each text to be processed is parsed through the Map operator (distributed computing program), and each text to be processed is segmented, and a stop word dictionary is loaded to remove stop words in the text to obtain the target text to be processed.

[0143] In step 12, each target text to be processed is segmented and N-Gram grouped.

[0144] Among them, the 2-Gram words of the same target text to be processed can be put into the word set b2, the 3-Gram words into the word set b3, and the 4-Gram words into the word set b4. Therefore, each target text to be processed is split into an N-Gram data set S i ={b2, b3, ..., bn}, and different N-gram sets have different weights, for example, the set weight of 2-gram is set to β 2 , the N-gram set weight is set to β n , different weight sets {β 2 , β 3 , ...β n}.

[0145] In step 13, the word vector is trained using the word2vec algorithm.

[0146] Among them, each word segmentation is trained through the word2vec algorithm and the vector of each word segmentation is calculated.

[0147] In step 14, the word frequency information of all word segments is calculated.

[0148] Among them, the word frequency p of different word segments w in each N-gram set is calculated w , and each word segment w and word frequency p w The mapping relationship is saved in the hdfs path.

[0149] In step 15, the words are weighted and summed to obtain the sentence vector.

[0150] The word vector dictionary is loaded from the hdfs path, and then the different word segments w in each N-gram set are mapped to the corresponding word vector v w , v w The length of is d.

[0151] Find the 2-gram word set b 2 The word vector g 2 .

[0152]

[0153] Among them, the g 2 is the word vector of the word segmentation set corresponding to any target text to be processed when the word length is 2. α is a hyperparameter, that is, a known parameter. w is the vector of each word segment in the hierarchical set with a word length of 2, w is the number of word segments, and the p w is the frequency of each word segment. Based on the above formula, the vector v of each word segment is w Divide by the frequency of the segmented word and sum them up, so that the segmented words with higher frequency have lower proportion, and the segmented words with lower frequency have higher proportion, which is in line with practical applications. Finally, we get the word vector that better expresses the meaning of each segmented word set.

[0154] Find the 2-gram word set b 3 The word vector g 3 .

[0155]

[0156] Among them, the g 3 is the word vector of the word segmentation set corresponding to any target text to be processed when the word length is 3. α is a hyperparameter, that is, a known parameter. w is the vector of each word segment in the hierarchical set with a word length of 3, w is the number of word segments, and the p w is the frequency of each word segment. Based on the above formula, the vector v of each word segment is w Divide by the frequency of the segmented word and sum them up, so that the higher the frequency of the segmented word, the lower the proportion of the segmented word, and the lower the frequency of the segmented word, the higher the proportion of the segmented word, which is in line with the actual application. Finally, each segmented word set is obtained to express the word meaning better. In this way, for each target text to be processed S i = {b 2 , b 3 , ..., b n}, S i The sentence vector is W i .

[0157]

[0158] The W i is the vector of the sentence, the number of i is equal to the number of target texts to be processed, g i is the word vector of the corresponding word segmentation set of any target text to be processed with word length i. i is the weight of the word segmentation set corresponding to the word length i, and each word vector g of the same target text to be processed is i and the corresponding weight θ iMultiply them and sum up the products of the same text to be processed to obtain the sentence vector W corresponding to each target text to be processed. i .

[0159] In step 16, the principal component vector is obtained through PCA (Principal Component Analysis).

[0160] Among them, each sentence vector W can be obtained. i , and transfer all the sentence vectors W i to the Reduce operator for aggregation to reduce the dimension, thereby obtaining the sentence vector matrix X, X = {W 1 , W 2 , W 3 ,..., W i}, the dimension of the matrix is (i, d), where n is the number of all texts and d is the size of each sentence vector.

[0161] Furthermore, perform principal component analysis on the sentence vector matrix X through the PCA tool to obtain the principal component matrix μ. The principal component matrix μ is an important component matrix of the matrix X after dimensionality reduction, and can also be understood as the common information contained in the matrix X.

[0162] In step 17, the principal component vector is corrected by the sentence vector.

[0163] Among them, for X = {W 1 , W 2 , W 3 , …, W i}, each sentence vector W i in it is subtracted by the principal component matrix μ.

[0164]

[0165] Among them, the is the target sentence vector matrix, μ is the principal component matrix, and μ T is the transpose matrix of the principal component matrix. Thus, through the above formula, calculate the difference between the product of each sentence vector w i in the sentence vector matrix and the principal component matrix μ, the transpose matrix μ T and the corresponding sentence w i , and subtract the principal component matrix μ to remove the principal component vector corresponding to the common information of all sentences, obtaining the calculated target sentence vector matrix The remaining target sentence vector matrix can better represent the difference between itself and other target sentence vector matrices. After the above transformation, the matrix X is converted to and

[0166] In step 18, the similarity of the sentence vectors is calculated to determine whether they are repeated, and the final result is filtered and output.

[0167] Among them, through the Map operator Split to obtain multiple target sentence vectors Each target sentence vector A corresponding text id may be identified, and the text id is associated with a corresponding target text to be processed (ie, a sentence).

[0168] In this way, any two target sentence vectors can be calculated The cosine similarity between the two target texts is calculated, and the target sentence vectors whose cosine similarity is less than the preset cosine threshold are deduplicated for the corresponding target text pairs to be processed, that is, any one of the two text IDs associated with the target sentence vector pair is directly filtered. Finally, all target texts to be processed with repeated text content are deduplicated. Since the generation of the target sentence vector is combined with semantics and the common information is removed through principal component analysis, the similarity calculation between the target sentence vectors is more accurate, and a more comprehensive deduplication effect is achieved.

[0169] In order to better implement the information processing method provided in the embodiment of the present application, the embodiment of the present application also provides a device based on the above information processing method. The meanings of the terms are the same as those in the above information processing method, and the specific implementation details can refer to the description in the method embodiment.

[0170] See also Figure 6 , Figure 6 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application, wherein the information processing device may include an acquisition unit 301, a word segmentation unit 302, a first calculation unit 303, a removal unit 304 and a second calculation unit 305, etc.

[0171] The acquisition unit 301 is used to acquire a target text set to be processed, where the target text set to be processed includes a plurality of target texts to be processed.

[0172] In some embodiments, the acquisition unit 301 is used to:

[0173] Obtain a text set to be processed, which contains multiple texts to be processed;

[0174] The stop words in each to-be-processed text are removed to obtain a plurality of target to-be-processed texts after the removal to generate a target to-be-processed text set.

[0175] The word segmentation unit 302 is used to perform word segmentation processing on each target text to be processed according to different word lengths in sequence, so as to obtain a word segmentation set corresponding to each word length of each target text to be processed.

[0176] In some embodiments, the word segmentation unit is used to:

[0177] Each target text to be processed is segmented in turn according to the sliding windows corresponding to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed.

[0178] The first calculation unit 303 is used to calculate the sentence vector corresponding to each target text to be processed based on the word vector corresponding to each word segmentation set.

[0179] In some embodiments, the first computing unit includes:

[0180] The first calculation subunit is used to sequentially calculate the word vector of each word segmentation set corresponding to each target text to be processed;

[0181] Set subunits to set different weights for each word vector according to the word length corresponding to each word segmentation set;

[0182] The second calculation subunit is used to calculate each word vector and the corresponding weight of the same target text to be processed to obtain the sentence vector corresponding to each target text to be processed.

[0183] In some embodiments, the first computing subunit is configured to:

[0184] Get the vector and word frequency information corresponding to each word in each word set;

[0185] The vector of each word in each word set in the same target text to be processed and the corresponding word frequency information are calculated to obtain the word vector of each word set corresponding to each target text to be processed.

[0186] In some embodiments, the second computing subunit is configured to:

[0187] Multiply each word vector of each target text to be processed by the corresponding weight to obtain multiple products corresponding to each target text to be processed;

[0188] Multiple products corresponding to the same target text to be processed are summed to obtain the sentence vector corresponding to each target text to be processed.

[0189] The removal unit 304 is used to remove the principal component vector in each sentence vector to obtain multiple target sentence vectors after the removal process.

[0190] In some embodiments, the removing unit 304 includes:

[0191] A combination subunit, used to combine each sentence vector to obtain a sentence vector matrix;

[0192] The analysis subunit is used to perform principal component analysis on the sentence vector matrix to obtain a principal component vector matrix;

[0193] The removal subunit is used to remove each sentence vector in the sentence vector matrix from the principal component vector matrix in turn to obtain a target sentence vector matrix, which includes multiple target sentence vectors.

[0194] In some embodiments, the removal subunit is used to:

[0195] Get the transposed matrix corresponding to the principal component vector matrix;

[0196] The difference between each sentence vector in the sentence vector matrix and the product of the principal component vector matrix, the transposed matrix and the corresponding sentence vector is calculated to obtain the calculated target sentence vector matrix.

[0197] The second calculation unit 305 is used to calculate the similarity between each target sentence vector, and perform deduplication processing on the target sentence vectors with similarity less than a preset threshold and the corresponding target text pairs to be processed.

[0198] In some embodiments, the second computing unit 305 is configured to:

[0199] Split the calculated target sentence vector matrix to obtain multiple target sentence vectors;

[0200] The cosine similarity between each target sentence vector is calculated, and the target sentence vectors whose cosine similarity is less than a preset cosine threshold are deduplicated to the corresponding target text pairs to be processed.

[0201] The specific implementation of each of the above units can be found in the previous embodiments, which will not be described in detail here.

[0202] As can be seen from the above, the embodiment of the present application obtains the target to-be-processed text set through the acquisition unit 301; the word segmentation unit 302 performs word segmentation processing on each target to-be-processed text according to different word lengths in turn, and obtains the word segmentation set corresponding to each word length of each target to-be-processed text; the first calculation unit 303 calculates the sentence vector corresponding to each target to-be-processed text based on the word vector corresponding to each word segmentation set; the removal unit 304 removes the principal component vector in each sentence vector, and obtains multiple target sentence vectors after the removal processing; the second calculation unit 305 calculates the similarity between each target sentence vector, and performs deduplication processing on the target sentence vector with a similarity less than a preset threshold value for the corresponding target to-be-processed text pair. In this way, a sentence vector with accurate semantic expression is generated through word segmentation processing, and the principal component vector of the sentence vector is removed, so that the difference between the vectors is more obvious, so that when judging the similarity of sentences, different sentences can be better distinguished. Compared with the existing manual text deduplication method, the embodiment of the present application can realize an automatic and accurate text deduplication method, which improves the efficiency and accuracy of information processing.

[0203] The present application also provides a server, such as Figure 7 As shown, it shows a schematic diagram of the structure of the server involved in the embodiment of the present application, specifically:

[0204] The server may include one or more processing core processors 401, one or more computer-readable storage media memories 402, a power supply 403, an input unit 404 and other components. Those skilled in the art will appreciate that Figure 7 The server structure shown in the figure does not constitute a limitation on the server, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:

[0205] The processor 401 is the control center of the server, and uses various interfaces and lines to connect various parts of the entire server. By running or executing software programs and / or modules stored in the memory 402, and calling data stored in the memory 402, the processor 401 performs various functions of the server and processes data, thereby monitoring the server as a whole. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 401.

[0206] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the server, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0207] The server also includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so that the power management system can manage charging, discharging, power consumption and other functions. The power supply 403 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators and other arbitrary components.

[0208] The server may further include an input unit 404, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0209] Although not shown, the server may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 401 in the server will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402, thereby realizing various functions, as follows:

[0210] A target text set to be processed is obtained, which contains multiple target texts to be processed; each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed; a sentence vector corresponding to each target text to be processed is calculated based on the word vector corresponding to each segmentation set; a principal component vector in each sentence vector is removed to obtain multiple target sentence vectors after the removal; the similarity between each target sentence vector is calculated, and the target sentence vector with a similarity less than a preset threshold is deduplicated to the corresponding target text to be processed.

[0211] In the above embodiments, the description of each embodiment has its own focus. For the part that is not described in detail in a certain embodiment, please refer to the detailed description of the information processing method above, and will not be repeated here.

[0212] As can be seen from the above, the server of the embodiment of the present application can obtain a target set of texts to be processed; perform word segmentation processing on each target text to be processed according to different word lengths in turn, and obtain a word segmentation set corresponding to each word length of each target text to be processed; calculate the sentence vector corresponding to each target text to be processed based on the word vector corresponding to each word segmentation set; remove the principal component vector in each sentence vector to obtain multiple target sentence vectors after removal; calculate the similarity between each target sentence vector, and perform deduplication processing on the corresponding target text to be processed for the target sentence vector whose similarity is less than a preset threshold. In this way, a sentence vector with accurate semantic expression is generated through word segmentation processing, and the principal component vector of the sentence vector is removed to make the difference between the vectors more obvious, so that different sentences can be better distinguished when judging the similarity of sentences. Compared with the existing manual text deduplication method, the embodiment of the present application can realize an automatic and accurate text deduplication method, which improves the efficiency and accuracy of information processing.

[0213] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0214] To this end, an embodiment of the present application provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions can be loaded by a processor to execute the steps in any one of the information processing methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:

[0215] A target text set to be processed is obtained, which contains multiple target texts to be processed; each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed; a sentence vector corresponding to each target text to be processed is calculated based on the word vector corresponding to each segmentation set; the principal component vector in each sentence vector is removed to obtain multiple target sentence vectors after the removal; the similarity between each target sentence vector is calculated, and the target sentence vector with a similarity less than a preset threshold is deduplicated to the corresponding target text to be processed.

[0216] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0217] Since the instructions stored in the computer-readable storage medium can execute the steps in any information processing method provided in the embodiments of the present application, the beneficial effects that can be achieved by any information processing method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0218] According to one aspect of an embodiment of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes a computer instruction, and the computer instruction is stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the method provided in various optional implementations provided in the above embodiments.

[0219] The above is a detailed introduction to an information processing method, device and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An information processing method, It is characterized in that include: Acquire a target text set to be processed, wherein the target text set to be processed includes a plurality of target texts to be processed; Each target text to be processed is segmented in turn according to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed; Calculate the sentence vector of each target text to be processed based on the word vector corresponding to each word segmentation set; The principal component vector in each sentence vector is removed to obtain multiple target sentence vectors after the removal process; The similarity between each target sentence vector is calculated, and the target sentence vectors with similarity less than a preset threshold are deduplicated to the corresponding target text pairs to be processed.

2. The information processing method according to claim 1, It is characterized in that The step of calculating the sentence vector corresponding to each target text to be processed based on the word vector corresponding to each word segmentation set includes: Calculate the word vector of each word segmentation set corresponding to each target text to be processed in turn; According to the word length corresponding to each word segmentation set, different weights are set for each word vector; Calculate each word vector and the corresponding weight of the same target text to be processed to obtain the corresponding sentence vector of each target text to be processed.

3. The information processing method according to claim 2, It is characterized in that The step of sequentially calculating the word vectors of each word segmentation set corresponding to each target text to be processed includes: Get the vector and word frequency information corresponding to each word in each word set; The vector of each word in each word set in the same target text to be processed and the corresponding word frequency information are calculated to obtain the word vector of each word set corresponding to each target text to be processed.

4. The information processing method according to claim 2, It is characterized in that The step of calculating each word vector and the corresponding weight of the same target text to be processed to obtain the sentence vector corresponding to each target text to be processed includes: Multiply each word vector of each target text to be processed by the corresponding weight to obtain multiple products corresponding to each target text to be processed; Multiple products corresponding to the same target text to be processed are summed to obtain the sentence vector corresponding to each target text to be processed.

5. The information processing method according to claim 1, It is characterized in that The principal component vector in each sentence vector is removed to obtain multiple target sentence vectors after the removal process, including: Combine each sentence vector to get the sentence vector matrix; Performing principal component analysis on the sentence vector matrix to obtain a principal component vector matrix; The principal component vector matrix is ​​removed from each sentence vector in the sentence vector matrix in turn to obtain a target sentence vector matrix, wherein the target sentence vector matrix includes multiple target sentence vectors.

6. The information processing method according to claim 5, It is characterized in that The step of removing the principal component vector matrix from each sentence vector in the sentence vector matrix in turn to obtain a target sentence vector matrix includes: Obtaining a transposed matrix corresponding to the principal component vector matrix; The difference between each sentence vector in the sentence vector matrix and the product of the principal component vector matrix, the transposed matrix and the corresponding sentence vector is calculated to obtain a calculated target sentence vector matrix.

7. The information processing method according to claim 6, It is characterized in that The calculating the similarity between each target sentence vector and performing deduplication processing on the target sentence vector with similarity less than a preset threshold and the corresponding target text pair to be processed includes: Split the calculated target sentence vector matrix to obtain multiple target sentence vectors; The cosine similarity between each target sentence vector is calculated, and the target sentence vectors whose cosine similarity is less than a preset cosine threshold are deduplicated to the corresponding target text pairs to be processed.

8. The information processing method according to any one of claims 1 to 7, It is characterized in that The step of obtaining a target text set to be processed includes: Obtaining a text set to be processed, wherein the text set to be processed includes a plurality of texts to be processed; The stop words in each to-be-processed text are removed to obtain a plurality of target to-be-processed texts after the removal to generate a target to-be-processed text set.

9. The information processing method according to any one of claims 1 to 7, It is characterized in that The word segmentation process is performed on each target text to be processed according to different word lengths in sequence to obtain a word segmentation set corresponding to each word length of each target text to be processed, including: Each target text to be processed is segmented in turn according to the sliding windows corresponding to different word lengths to obtain a segmentation set corresponding to each word length of each target text to be processed.

10. An information processing device, It is characterized in that include: An acquisition unit, used for acquiring a target text set to be processed, wherein the target text set to be processed includes a plurality of target texts to be processed; The word segmentation unit is used to perform word segmentation processing on each target text to be processed according to different word lengths in turn, and obtain a word segmentation set corresponding to each word length of each target text to be processed; A first calculation unit is used to calculate a sentence vector corresponding to each target text to be processed based on the word vector corresponding to each word segmentation set; A removal unit, used for removing the principal component vector in each sentence vector to obtain a plurality of target sentence vectors after the removal process; The second calculation unit is used to calculate the similarity between each target sentence vector, and perform deduplication processing on the target sentence vector with a similarity less than a preset threshold and the corresponding target text to be processed.

11. The information processing device according to claim 10, It is characterized in that The first computing unit is configured to: The first calculation subunit is used to sequentially calculate the word vector of each word segmentation set corresponding to each target text to be processed; Set subunits to set different weights for each word vector according to the word length corresponding to each word segmentation set; The second calculation subunit is used to calculate each word vector and the corresponding weight of the same target text to be processed to obtain the sentence vector corresponding to each target text to be processed.

12. The information processing device according to claim 11, It is characterized in that The first computing subunit is configured to: Get the vector and word frequency information corresponding to each word in each word set; The vector of each word in each word set in the same target text to be processed and the corresponding word frequency information are calculated to obtain the word vector of each word set corresponding to each target text to be processed.

13. The information processing device according to claim 11, It is characterized in that The second computing subunit is configured to: Multiply each word vector of each target text to be processed by the corresponding weight to obtain multiple products corresponding to each target text to be processed; Multiple products corresponding to the same target text to be processed are summed to obtain the sentence vector corresponding to each target text to be processed.

14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the information processing method according to any one of claims 1 to 9.

15. A computer program product comprising a computer program or instructions, It is characterized in that When the computer program or instruction is executed by a processor, the steps of the information processing method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Text aggregation and display method and system in large-scale data-oriented intelligence system

    CN106294861A

  • Mixed multi-feature sentence similarity calculation method and system, and storage medium

    CN110705612A