Text similarity recognition method, system and equipment and storage medium

By performing word segmentation and embedding vector processing on the text, and calculating weighted similarity with attention matrix, the problems of cumbersome feature engineering and limited generalization capabilities in the existing technology are solved, and more efficient text similarity recognition is achieved.

CN120012777APending Publication Date: 2025-05-16SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411832199.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art requires a lot of feature engineering in text similarity recognition, and the generalization ability of the model is limited, making it difficult to effectively understand the context semantic information of the text.

Method used

By performing word segmentation processing on the text and converting it into integer indexes, the embedding matrix of the pretrained semantic recognition model converts the integer index into an embedding vector, calculates the Euclidean distance of different sets of vectors, and calculates the weighted similarity based on the Euclidean distance using the attention matrix.

Benefits of technology

It improves the accuracy of text similarity recognition, can more effectively understand the context semantic information and deep meaning of the text, and has reliable design principles, simple structure, and has broad application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012777A_ABST
    Figure CN120012777A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and particularly provides a text similarity recognition method, system and device and a storage medium, and the method comprises the steps: carrying out word segmentation processing on a text, and converting a text unit after word segmentation into an integer index; converting each integer index into a corresponding embedding vector through an embedding matrix of a pre-training semantic recognition model, and storing the embedding vectors of the same text into the same vector set; euclidean distances of vectors of different vector sets are calculated, and weighted similarity is calculated based on the Euclidean distances by using an attention matrix. According to the method, the text is converted into the high-dimensional vector, the context semantic information and the deep meaning of the text are fully understood, and then the similarity between the vectors is calculated through the attention weight, so that the accuracy of text similarity recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a text similarity recognition method, system, device and storage medium. Background Art

[0002] In the field of natural language processing (NLP), text similarity recognition has always been one of the research hotspots and challenges. In the early days, traditional text similarity recognition methods mainly included vocabulary-based methods and rule-based methods. Vocabulary-based methods, such as cosine similarity, evaluate similarity by calculating the angle between word frequency vectors in the text; while rule-based methods rely on manually formulated grammatical and semantic rules to compare texts. With the development of machine learning technology, some machine learning-based text similarity recognition methods have emerged, such as using support vector machines (SVM), decision trees, random forests and other algorithms to classify or predict text similarity. These methods usually require a lot of feature engineering to extract text features, and the generalization ability of the model is limited. Summary of the invention In view of the above-mentioned deficiencies in the prior art, the present invention provides a text similarity recognition method, system, device and storage medium to solve the above-mentioned technical problems.

[0003] In a first aspect, the present invention provides a text similarity recognition method, comprising: Perform word segmentation on the text and convert the text units after word segmentation into integer indexes; Through the embedding matrix of the pre-trained semantic recognition model, each integer index is converted into a corresponding embedding vector, and the embedding vectors of the same text are saved in the same vector set; The Euclidean distances of vectors of different vector sets are calculated, and the weighted similarity is calculated based on the Euclidean distance using an attention matrix.

[0004] In an optional implementation, the text is segmented and the segmented text units are converted into integer indexes, including: Clean the text to remove punctuation, numbers, and stop words; Segment text into words or phrases; Collect all the non-repeating words or phrases to form a vocabulary list; Convert tokenized text into integer indices corresponding to words or phrases in the vocabulary.

[0005] In an optional embodiment, each integer index is converted into a corresponding embedding vector through an embedding matrix of a pre-trained semantic recognition model, and the embedding vectors of the same text are saved in the same vector set, including: Add special tokens for integer indices and perform attention masking; Input the integer index into the pre-trained BERT model and obtain the last hidden state as the initial vector representation of the text; Initialize the pseudo token vector of batch size, initialize the pseudo space embedding matrix, and construct the mapping space; Mapping the initial vector representation into a mapping space; Add positional coding; The vector representation in the mapping space is reflected back, and the information in the mapping space, the information of the initial vector representation and the position encoding information are fused to obtain the final text vector representation.

[0006] In an optional implementation, calculating the Euclidean distance of vectors of different vector sets, and calculating the weighted similarity based on the Euclidean distance using an attention matrix, includes: Calculate the Euclidean distance between each vector in the first vector set and each vector in the second vector set to obtain a distance matrix; Normalizing the distance matrix; Initialize the attention matrix and use the negative exponential function to calculate the weight element value of the attention matrix; Converting the distance matrix into a similarity matrix; By calculating the product of the similarity matrix and the attention matrix, a new similarity matrix is ​​formed.

[0007] In a second aspect, the present invention provides a text similarity recognition system, comprising: The text processing module is used to perform word segmentation on the text and convert the text units after word segmentation into integer indexes; The semantic analysis module is used to convert each integer index into a corresponding embedding vector through the embedding matrix of the pre-trained semantic recognition model, and save the embedding vectors of the same text into the same vector set; The similarity calculation module is used to calculate the Euclidean distance of vectors of different vector sets, and calculate the weighted similarity based on the Euclidean distance using the attention matrix.

[0008] In an optional implementation, the text processing module includes: The preprocessing unit is used to clean the text and remove punctuation marks, numbers, and stop words; A text segmentation unit, used to segment text into words or phrases; A list generation unit is used to collect all non-repeated words or phrases to form a vocabulary list; The index generation unit is used to convert the tokenized text into integer indices corresponding to the words or phrases in the vocabulary.

[0009] In an optional embodiment, the semantic recognition module includes: Index processing unit, used to add special tokens to integer indices and perform attention masking; The initial mapping unit is used to input the integer index into the pre-trained BERT model and obtain the last hidden state as the initial vector representation of the text; The space construction unit is used to initialize the pseudo-token vector of the batch size, initialize the pseudo-space embedding matrix, and construct the mapping space; A secondary mapping unit, used for mapping the initial vector representation into a mapping space; Position coding unit, used to add position coding; The vector echo unit is used to echo the vector representation in the mapping space, fuse the information in the mapping space, the information of the initial vector representation and the position encoding information, and obtain the final text vector representation.

[0010] In an optional implementation, the similarity calculation module includes: A distance calculation unit, used to calculate the Euclidean distance between each vector in the first vector set and each vector in the second vector set to obtain a distance matrix; A normalization unit, used for normalizing the distance matrix; A weight calculation unit, used to initialize the attention matrix and calculate the weight element value of the attention matrix using a negative exponential function; A distance conversion unit, used to convert the distance matrix into a similarity matrix; The weighted calculation unit is used to form a new similarity matrix by calculating the product of the similarity matrix and the attention matrix.

[0011] In a third aspect, a device is provided, including: A memory, used for storing a text similarity recognition program; A processor is used to implement the steps of the text similarity recognition method provided in the first aspect when executing the text similarity recognition program.

[0012] In a fourth aspect, a computer-readable storage medium is provided, on which a text similarity recognition program is stored. When the text similarity recognition program is executed by a processor, the steps of the text similarity recognition method provided in the first aspect are implemented.

[0013] The beneficial effect of the present invention lies in that the text similarity recognition method, system, device and storage medium provided by the present invention convert text into high-dimensional vectors, fully understand the contextual semantic information and deep meaning of the text, and then calculate the similarity between vectors through attention weights, thereby improving the accuracy of text similarity recognition.

[0014] In addition, the invention has a reliable design principle, a simple structure and a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.

[0017] Figure 2 is a schematic block diagram of a system according to an embodiment of the present invention.

[0018] Figure 3 A schematic diagram of the structure of a device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0021] The key terms appearing in the present invention are explained below.

[0022] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language representation model developed by the Google AI Language team in 2018. BERT is based on the Transformer architecture, especially its encoder part, and aims to capture the deep features of language through large-scale unsupervised training, thereby achieving excellent performance in various natural language processing (NLP) tasks.

[0023] The text similarity recognition method provided by the embodiment of the present invention is executed by a computer device, and accordingly, the text similarity recognition system runs in the computer device.

[0024] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention. Figure 1 The execution subject may be a text similarity recognition system. According to different requirements, the order of the steps in the flow chart may be changed, and some may be omitted.

[0025] like Figure 1 As shown, the method includes: S1. Segment the text and convert the segmented text units into integer indices.

[0026] In natural language processing (NLP) tasks, text first needs to be converted into a form that computers can understand and process. This step usually includes word segmentation and mapping the segmented text units (such as words or subwords) to integer indices.

[0027] Word segmentation: Word segmentation is the process of dividing a continuous text string into smaller, more meaningful units. These units can be words, subwords (such as a combination of roots and affixes), characters, or byte pair encoding (BPE) units. The word segmentation method depends on the specific NLP task and the language used.

[0028] Convert to integer index: Once the text is tokenized, each token unit needs to be mapped to a unique integer index. This is usually done through a predefined vocabulary that contains all possible token units and their corresponding integer indexes. This process is also called the initial preparation of text vectorization or word embedding.

[0029] S2. Convert each integer index into a corresponding embedding vector through the embedding matrix of the pre-trained semantic recognition model, and save the embedding vectors of the same text into the same vector set.

[0030] After obtaining the integer index representation of the text, the next step is to convert these indices into richer embedding vectors that contain semantic information.

[0031] Pre-trained semantic recognition model: This refers to pre-trained language models such as BERT, which can generate high-quality word embeddings by learning the statistical laws of language on large-scale text data.

[0032] Embedding matrix: Pre-trained models usually contain an embedding matrix, which is a weight matrix where each row corresponds to the embedding vector of a word unit in the vocabulary. The corresponding embedding vector can be retrieved by using an integer index as the row index of the embedding matrix.

[0033] Save to vector set: For each word segmentation unit in the text, convert its integer index into an embedded vector and save these vectors in the same vector set. In this way, an embedded vector representation of the text is obtained, which retains the semantic information of the text.

[0034] S3. Calculate the Euclidean distance of vectors of different vector sets, and calculate the weighted similarity based on the Euclidean distance using the attention matrix.

[0035] After obtaining the embedded vector representation of the text, the similarity between different texts can be calculated.

[0036] Calculate Euclidean distance: Euclidean distance is a commonly used distance metric used to calculate the straight-line distance between two points (here, embedded vectors). For the embedded vector sets of two texts A and B, the Euclidean distance of all vector pairs between them can be calculated to obtain a distance matrix.

[0037] Attention Matrix: Attention mechanism is a method for calculating the relevance between different parts. Here, the attention matrix can be used to calculate the weighted similarity based on the Euclidean distance. Specifically, the Euclidean distance can be converted to a similarity score (for example, using a negative exponential function), and then an attention matrix is ​​constructed, where each element represents the weighted similarity between two embedding vectors.

[0038] Weighted Similarity: Through the attention matrix, we can get the weighted similarity between different texts. This similarity score can be used in various NLP tasks such as text classification, information retrieval, text summarization, etc.

[0039] In an embodiment of the present invention, based on step S1, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.

[0040] S101. Text cleaning Text cleaning is the first step in text preprocessing, aiming to remove irrelevant information in the text, such as punctuation marks, numbers, stop words, etc., to simplify the text and reduce noise.

[0041] Removing punctuation marks: Punctuation marks are usually used to separate sentences, phrases or words, but in NLP tasks, they often do not contain useful information. Therefore, all punctuation marks in the text need to be removed.

[0042] Removing numbers: Numbers in the text may represent dates, times, quantities, etc., but in some NLP tasks, such as sentiment analysis, text classification, etc., the specific values of numbers are not important. Therefore, to simplify the text, all numbers can be chosen to be removed.

[0043] Removing stop words: Stop words are words that frequently appear in a language but contribute little to the meaning of the text, such as "de", "le", "zai", etc. Removing stop words can reduce the redundancy of the text and improve the efficiency of subsequent processing.

[0044] S102. Word segmentation Word segmentation is the process of splitting a continuous text string into smaller and more meaningful units. These units can be words, phrases or sub-words, etc.

[0045] Rule-based word segmentation: Use predefined rules and dictionaries for word segmentation, such as forward maximum matching, backward maximum matching, etc. This method is simple and effective, but depends on the integrity and accuracy of the dictionary.

[0046] Statistics-based word segmentation: Use machine learning algorithms to segment the text, such as Hidden Markov Model (HMM), Conditional Random Field (CRF), etc. This method can automatically learn the rules of word segmentation, but requires training data and computing resources.

[0047] Deep learning-based word segmentation: Use deep learning models (such as BERT, GPT, etc.) for word segmentation. This method can capture more complex language features, but requires more computing resources and training time.

[0048] S103. Building a vocabulary A vocabulary is a list that contains all non-repeating words or phrases, and is used to convert the segmented text into integer indices.

[0049] Collecting non-repeating words or phrases: Traverse the segmented text and collect all non-repeating words or phrases to form a vocabulary.

[0050] Sorting and indexing: Sort the words or phrases in the vocabulary and assign a unique integer index to each word or phrase. In this way, the segmented text can be converted into a sequence of integer indices.

[0051] In an embodiment of the present invention, based on step S2, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.

[0052] BERT pre-training tasks: Masked Language Model (MLM): Randomly mask some tokens in the input text and let the model predict these masked tokens. Next Sentence Prediction (NSP): Determine whether two sentences are continuous contexts.

[0053] The method for identifying text semantic vectors based on the BERT model is as follows: S201: Add special token and attention mask processing for integer indexes Special token addition: When converting text to integer indexes, in addition to the regular vocabulary index, you also need to add some special tokens, such as CLS (used for classification tasks, usually located at the beginning of the sequence) and SEP (used to separate sentences, such as sentence separation when there are multiple sentences input).

[0054] These special tokens have specific roles in the BERT model. For example, the output of the CLS token is usually used for the final prediction of the classification task.

[0055] Attention Mask Processing: The attention mask is used to distinguish between actual input tokens and padding tokens to ensure that the model ignores the padding tokens when calculating attention.

[0056] Create a mask sequence of the same length as the input sequence, where the actual input token position is 1 and the padding token position is 0.

[0057] Attention masks are particularly important when dealing with variable-length inputs, allowing the model to focus only on the valid parts of the input while ignoring the additional noise due to padding.

[0058] S202: Integer index input pre-trained BERT model to obtain initial vector representation Model input preparation: A sequence of integer indices containing special tokens and attention masks is given as input to the BERT model.

[0059] At the same time, it is also necessary to prepare a segmentation mask (used to distinguish different sentences when multiple sentences are input), but it may not be required in some single sentences or specific tasks.

[0060] Get the initial vector representation: Feed the preprocessed input sequence into the pre-trained BERT model.

[0061] Perform forward propagation through the Transformer layer of the model and extract the last hidden state as the initial vector representation of the text.

[0062] These representations usually have a fixed dimension (such as 768 dimensions for BERT-base) and contain semantic information of the text.

[0063] S203: Initialize pseudo-token vector, embedding matrix and construct mapping space Initialize the pseudo token vector: For each example in the batch, initialize a vector with all specific values ​​(such as 0 or the index corresponding to [MASK]), or you can initialize it to a zero vector.

[0064] The dimensions of these vectors are different from the dimensions of the BERT output vectors (e.g. 128 dimensions), with the goal of creating a pseudo-space that is different from the BERT output space.

[0065] The embedding matrix E is initialized: Initialize an embedding matrix E whose dimensions depend on the dimensions of the pseudo-token vector and the BERT output vector.

[0066] This matrix is ​​used to map the pseudo-token vector to the same dimensional space as the BERT output vector for subsequent self-attention calculations.

[0067] Mapping space construction: Multiply each pseudo-token vector by the embedding matrix E to obtain a set of vectors equal to the batch size.

[0068] These vectors constitute the mapping space, which is initially semantically non-semantic but will be used in subsequent self-attention calculations to fuse the original information generated by BERT with the information of the mapping space.

[0069] S204: Mapping the initial vector representation into the mapping space Self-Attention module initialization: Initialize a Self-Attention module, which includes linear transformations of query, key, and value, and a softmax function to calculate the attention weight.

[0070] Self-attention calculation: Use the initial vector representation generated by BERT as Key and Value.

[0071] The vector in the mapping space is used as the query.

[0072] The updated vector representations are calculated through the Self-Attention module. These representations are now in the mapping space and combine the original information generated by BERT and the information of the mapping space.

[0073] S205: Position coding added Position encoding function: Since the BERT model itself is based on the Transformer architecture, it does not have a natural perception of the token position information in the sequence.

[0074] Therefore, adding positional encodings in the input sequence is necessary to provide position information so that the model can distinguish tokens at different positions.

[0075] Positional encoding method: Common methods include positional encoding of sine and cosine functions, which is widely used in Transformer models.

[0076] The dimensions of the positional encodings typically match the dimensions of the BERT output vectors so that they can be directly added or concatenated.

[0077] S206: Reflect the vector representation in the mapping space to obtain the final text vector representation Another Self-Attention module is initialized: Initialize another Self-Attention module that does not share parameters with the module in S204 to allow the model to learn different information under two different attention mechanisms.

[0078] Echo calculation: The vector representation in the mapping space obtained in S204 is used as Key and Value.

[0079] The initial vector representation (or the representation after some transformation) generated by BERT is used as the query.

[0080] The second Self-Attention module calculates the reflected vector representations. These representations fuse the information in the mapping space with the information of the original BERT representation (and possible positional encoding information), providing a richer text representation.

[0081] In summary, by adding special tokens, processing attention masks, initializing pseudo-token vectors and embedding matrices, building mapping spaces, performing self-attention calculations, and adding position encoding, we can obtain a text vector representation that contains both the original semantic information generated by the BERT model and the mapping space information. This representation method may have higher performance and flexibility when processing complex text tasks.

[0082] In an embodiment of the present invention, based on step S3, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.

[0083] S301. Vector collection preparation: Suppose there are two vector sets, vector set A (from text A) and vector set B (from text B), where A contains m vectors and B contains n vectors.

[0084] Euclidean distance calculation: For each vector a in the set A i (i=1,2,...,m) and each vector b in set B j (j=1,2,...,n), calculate the Euclidean distance between them.

[0085] Get an m×n distance matrix D, where D ij Represents vector a i and vector b j The Euclidean distance between .

[0086] Distance Matrix Processing: The distance matrix D can be normalized for better comparison and analysis. Normalization can be achieved by dividing each distance value by the maximum value in the distance matrix, or using other normalization methods.

[0087] S302: Calculate weighted similarity based on Euclidean distance using attention matrix Initialization of attention matrix: Initialize an attention matrix W with dimension m×n.

[0088] This matrix will be used to store weighted similarities based on Euclidean distance.

[0089] Attention weight calculation: For each element Dij in the distance matrix D, an attention weight wij is calculated.

[0090] The weight element values ​​of the attention matrix A are calculated using a negative exponential function (such as the softmax function). Specifically, for each element A[i][j] in the attention matrix A, its weight value A_weight[i][j] = exp(-D_norm[i][j] / temperature) is calculated, where temperature is a hyperparameter used to control the smoothness of the weight distribution. Then, the weight values ​​of each row (or column, depending on the specific task) are normalized so that their sum is 1.

[0091] S303. Convert the distance matrix into a similarity matrix Similarity matrix definition: The similarity matrix is ​​the inverse of the distance matrix, that is, the smaller the distance, the higher the similarity. Therefore, each element in the distance matrix D_norm can be inverted (or other transformations such as 1 / (1 + D_norm[i][j])) to obtain the similarity matrix S.

[0092] Forming the similarity matrix: According to the above conversion method, the distance matrix D_norm is converted into a similarity matrix S. The element S[i][j] in the similarity matrix S represents the similarity between the vector a_i and the vector b_j.

[0093] S304. Calculate a new similarity matrix Matrix multiplication: A new similarity matrix S_new is formed by calculating the product of the similarity matrix S and the attention matrix A_weight. Specifically, for each element S_new[i][k] in the new similarity matrix S_new, it is equal to the dot product of the i-th row of the similarity matrix S and the k-th column of the attention matrix A_weight (i.e. sum(S[i][j]* A_weight[j][k]), where j is the index of all columns).

[0094] Result interpretation: The new similarity matrix S_new combines the information of the original similarity and attention weights, and can be used for subsequent weighted summation, classification, clustering and other tasks.

[0095] It should be noted that the specific implementation details of the attention weight calculation and weighted similarity calculation in the above steps may need to be adjusted and optimized according to the actual tasks and datasets.

[0096] In some embodiments, the text similarity recognition system may include a plurality of functional modules composed of computer program segments. The computer programs of the various program segments in the text similarity recognition system may be stored in a memory of a computer device and executed by at least one processor to perform (see Figure 1 Description) The function of text similarity recognition.

[0097] In this embodiment, the text similarity recognition system can be divided into multiple functional modules according to the functions it performs, such as Figure 2 As shown. The functional modules of the system 200 may include: a text processing module 210, a semantic analysis module 220 and a similarity calculation module 230. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, which are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0098] The text processing module is used to perform word segmentation on the text and convert the text units after word segmentation into integer indexes; The semantic analysis module is used to convert each integer index into a corresponding embedding vector through the embedding matrix of the pre-trained semantic recognition model, and save the embedding vectors of the same text into the same vector set; The similarity calculation module is used to calculate the Euclidean distance of vectors of different vector sets, and calculate the weighted similarity based on the Euclidean distance using the attention matrix.

[0099] Optionally, as an embodiment of the present invention, the text processing module includes: The preprocessing unit is used to clean the text and remove punctuation marks, numbers, and stop words; A text segmentation unit, used to segment text into words or phrases; A list generation unit is used to collect all non-repeated words or phrases to form a vocabulary list; The index generation unit is used to convert the tokenized text into integer indices corresponding to the words or phrases in the vocabulary.

[0100] Optionally, as an embodiment of the present invention, the semantic recognition module includes: Index processing unit, used to add special tokens to integer indices and perform attention masking; The initial mapping unit is used to input the integer index into the pre-trained BERT model and obtain the last hidden state as the initial vector representation of the text; The space construction unit is used to initialize the pseudo-token vector of the batch size, initialize the pseudo-space embedding matrix, and construct the mapping space; A secondary mapping unit, used for mapping the initial vector representation into a mapping space; Position coding unit, used to add position coding; The vector echo unit is used to echo the vector representation in the mapping space, fuse the information in the mapping space, the information of the initial vector representation and the position encoding information, and obtain the final text vector representation.

[0101] Optionally, as an embodiment of the present invention, the similarity calculation module includes: A distance calculation unit, used to calculate the Euclidean distance between each vector in the first vector set and each vector in the second vector set to obtain a distance matrix; A normalization unit, used for normalizing the distance matrix; A weight calculation unit, used to initialize the attention matrix and calculate the weight element value of the attention matrix using a negative exponential function; A distance conversion unit, used to convert the distance matrix into a similarity matrix; The weighted calculation unit is used to form a new similarity matrix by calculating the product of the similarity matrix and the attention matrix.

[0102] Figure 3 The text similarity recognition method provided for the embodiment of the present application can be applied to equipment. It will be appreciated by those skilled in the art that the equipment structure involved in the embodiment of the present invention does not constitute a limitation on the equipment, and the equipment may include more or less components than shown, or combine certain components, or arrange different components. In an embodiment of the present invention, the equipment includes but is not limited to laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The equipment may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.

[0103] The device 300 may include: a processor 310, a memory 320 and a communication unit 330. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention, and it may be a bus structure or a star structure, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0104] The memory 320 may be used to store the execution instructions of the processor 310, and the memory 320 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the device 300 is enabled to perform some or all of the steps in the following method embodiments.

[0105] The processor 310 is the control center of the storage device, and uses various interfaces and lines to connect various parts of the entire electronic device. It runs or executes software programs and / or modules stored in the memory 320, and calls data stored in the memory to perform various functions of the electronic device and / or process data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of a plurality of packaged ICs with the same or different functions. For example, the processor 310 can include only a central processing unit (CPU). In an embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.

[0106] The communication unit 330 is used to establish a communication channel so that the storage device can communicate with other devices, receive user data sent by other devices or send user data to other devices.

[0107] The present invention also provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, the program may include some or all of the steps in each embodiment provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).

[0108] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes, including several instructions for enabling a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.

[0109] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.

[0110] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or modules, which can be electrical, mechanical or other forms.

[0111] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0112] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0113] Although the present invention has been described in detail with reference to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention. Any person of ordinary skill in the art may easily think of changes or substitutions within the technical scope disclosed by the present invention, and these shall be within the scope of protection of the present invention.

Claims

1. A text similarity recognition method, characterized in that: include: Perform word segmentation on the text and convert the text units after word segmentation into integer indexes; Through the embedding matrix of the pre-trained semantic recognition model, each integer index is converted into a corresponding embedding vector, and the embedding vectors of the same text are saved in the same vector set; The Euclidean distances of vectors of different vector sets are calculated, and the weighted similarity is calculated based on the Euclidean distance using an attention matrix.

2. The method according to claim 1, characterized in that: Perform word segmentation on the text and convert the word-segmented text units into integer indexes, including: Clean the text to remove punctuation, numbers, and stop words; Segment text into words or phrases; Collect all the non-repeating words or phrases to form a vocabulary list; Convert tokenized text into integer indices corresponding to words or phrases in the vocabulary.

3. The method according to claim 1, characterized in that Through the embedding matrix of the pre-trained semantic recognition model, each integer index is converted into a corresponding embedding vector, and the embedding vectors of the same text are saved to the same vector set, including: Add special tokens for integer indices and perform attention masking; Input the integer index into the pre-trained BERT model and obtain the last hidden state as the initial vector representation of the text; Initialize the pseudo token vector of batch size, initialize the pseudo space embedding matrix, and construct the mapping space; Mapping the initial vector representation into a mapping space; Add positional coding; The vector representation in the mapping space is reflected back, and the information in the mapping space, the information of the initial vector representation and the position encoding information are fused to obtain the final text vector representation.

4. The method according to claim 1, characterized in that: Calculating the Euclidean distance of vectors of different vector sets, and calculating the weighted similarity based on the Euclidean distance using the attention matrix, including: Calculate the Euclidean distance between each vector in the first vector set and each vector in the second vector set to obtain a distance matrix; Normalizing the distance matrix; Initialize the attention matrix and use the negative exponential function to calculate the weight element value of the attention matrix; Converting the distance matrix into a similarity matrix; By calculating the product of the similarity matrix and the attention matrix, a new similarity matrix is ​​formed.

5. A text similarity recognition system, characterized in that: include: The text processing module is used to perform word segmentation on the text and convert the text units after word segmentation into integer indexes; The semantic analysis module is used to convert each integer index into a corresponding embedding vector through the embedding matrix of the pre-trained semantic recognition model, and save the embedding vectors of the same text into the same vector set; The similarity calculation module is used to calculate the Euclidean distance of vectors of different vector sets, and calculate the weighted similarity based on the Euclidean distance using the attention matrix.

6. The system according to claim 5, characterized in that The text processing module comprises: The preprocessing unit is used to clean the text and remove punctuation marks, numbers, and stop words; A text segmentation unit, used to segment text into words or phrases; A list generation unit is used to collect all non-repeated words or phrases to form a vocabulary list; The index generation unit is used to convert the tokenized text into integer indices corresponding to the words or phrases in the vocabulary.

7. The system according to claim 5, characterized in that The semantic recognition module comprises: Index processing unit, used to add special tokens to integer indices and perform attention masking; The initial mapping unit is used to input the integer index into the pre-trained BERT model and obtain the last hidden state as the initial vector representation of the text; The space construction unit is used to initialize the pseudo-token vector of the batch size, initialize the pseudo-space embedding matrix, and construct the mapping space; A secondary mapping unit, used for mapping the initial vector representation into a mapping space; Position coding unit, used to add position coding; The vector echo unit is used to echo the vector representation in the mapping space, fuse the information in the mapping space, the information of the initial vector representation and the position encoding information, and obtain the final text vector representation.

8. The system according to claim 5, characterized in that The similarity calculation module comprises: A distance calculation unit, used to calculate the Euclidean distance between each vector in the first vector set and each vector in the second vector set to obtain a distance matrix; A normalization unit, used for normalizing the distance matrix; A weight calculation unit, used to initialize the attention matrix and calculate the weight element value of the attention matrix using a negative exponential function; A distance conversion unit, used to convert the distance matrix into a similarity matrix; The weighted calculation unit is used to form a new similarity matrix by calculating the product of the similarity matrix and the attention matrix.

9. A device, characterized in that: include: A memory, used for storing a text similarity recognition program; A processor is used to implement the steps of the text similarity recognition method according to any one of claims 1 to 4 when executing the text similarity recognition program.

10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a text similarity recognition program, and when the text similarity recognition program is executed by a processor, the steps of the text similarity recognition method according to any one of claims 1 to 4 are implemented.

Citation Information

Cited By

  • Text classification method, electronic equipment and storage medium

    CN120744127A

  • Detection method, system and equipment for structured sensitive data and medium

    CN120805199A