A text processing method, device, computer device, and storage medium
By fusion of word vector matrix and decomposition of singular value in text processing methods, the problem of inability to accurately reflect text correlation in the prior art is solved, and a more accurate text correlation representation is achieved.
Patent Information
- Application Number
- CN202210790159.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-07-06
AI Technical Summary
In the prior art, the overall correlation of different texts is determined through compressed single word vectors, and the overall correlation of each text cannot be accurately reflected.
The word vector matrix of the first text and the second text is fused, and the singular value decomposition is performed to obtain a singular value matrix, and the correlation of the text in multiple vector dimensions is determined based on the singular value matrix.
Through the correlation degree under multiple matrix dimensions, the overall correlation between the first text and the second text is more accurately and comprehensively represented.
Smart Images

Figure CN115062626B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of natural language processing technology, and in particular, to a text processing method, apparatus, computer device, and storage medium. Background Art
[0002] Natural language processing is an important direction in the field of computer science. In some natural language processing scenarios, the keywords in different texts are respectively transformed into word vector matrices, and then the word vector matrices of each text are respectively compressed into a single word vector. By using the correlation relationship between the compressed word vectors, the overall relevance of different texts can be determined.
[0003] In the above processing process, the overall relevance of different texts is usually determined by a single compressed word vector, which cannot accurately reflect the overall relevance of each text. Summary of the Invention
[0004] Embodiments of the present disclosure at least provide a text processing method, apparatus, computer device, and storage medium.
[0005] In a first aspect, an embodiment of the present disclosure provides a text processing method, including:
[0006] Obtain a first word vector matrix corresponding to a first keyword in a first text and a second word vector matrix corresponding to a second keyword in a second text; wherein, the word vector matrix includes: word vectors respectively corresponding to a plurality of keywords; each of the word vectors includes: vector elements respectively corresponding to a plurality of vector dimensions;
[0007] Fuse the first word vector matrix and the second word vector matrix to obtain a fusion matrix, and perform singular value decomposition on the fusion matrix to obtain a singular value matrix;
[0008] Determine the relevance of the first text and the second text respectively in a plurality of the vector dimensions based on the singular value matrix.
[0009] In an optional embodiment, when performing singular value decomposition on the fusion matrix, a first target decomposition matrix and a second target decomposition matrix are further obtained; the first target decomposition matrix is used to represent the weights respectively corresponding to the semantics of the first text in a plurality of the vector dimensions; the second target decomposition matrix is used to represent the weights respectively corresponding to the semantics of the second text in a plurality of the vector dimensions;
[0010] The determining the relevance of the first text and the second text respectively in a plurality of the vector dimensions based on the singular value matrix includes:
[0011] Based on the first word vector matrix and the first target decomposition matrix, determine the first semantic keywords of the first text in multiple vector dimensions respectively, and based on the first word vector matrix and the second target decomposition matrix, determine the second semantic keywords of the second text in multiple vector dimensions respectively;
[0012] Based on the singular value matrix, determine the correlation degrees of the first semantic keywords and the second semantic keywords in multiple vector dimensions respectively.
[0013] In an optional implementation manner, the determining the first semantic keywords of the first text in multiple vector dimensions respectively based on the first word vector matrix and the first target decomposition matrix includes:
[0014] Based on the first word vector matrix and the first target decomposition matrix, determine the word vector compression matrix corresponding to the first word vector matrix;
[0015] Based on each word vector in the word vector compression matrix and the word vectors corresponding to multiple candidate words respectively, determine the first semantic keywords corresponding to each word vector.
[0016] In an optional implementation manner, the determining the word vector compression matrix corresponding to the first word vector matrix based on the first word vector matrix and the first target decomposition matrix includes:
[0017] Compress the first target decomposition matrix according to a preset number of singular values to obtain the compressed decomposition matrix corresponding to the first target decomposition matrix; wherein, the singular value is the value of the matrix element corresponding to the target vector dimension in the singular value matrix;
[0018] Based on the compressed decomposition matrix and the first word vector matrix, determine the word vector compression matrix corresponding to the first word vector matrix.
[0019] In an optional implementation manner, the obtaining the word vector matrix corresponding to the keyword in the text includes:
[0020] Perform word segmentation on the text to obtain the initial words included in the text;
[0021] Perform screening processing on the initial words according to a preset screening rule to obtain the keywords;
[0022] Based on the number of times each keyword appears in the text respectively and the word vectors corresponding to each keyword, construct the word vector matrix corresponding to the text;
[0023] Among them, the text includes a first text, the keyword includes a first keyword, and the word vector matrix includes the first word vector matrix;
[0024] Or,
[0025] the text includes a second text, the keyword includes a second keyword, and the word vector matrix includes the second word vector matrix.
[0026] In an alternative embodiment, constructing the word vector matrix corresponding to the text based on the number of times each keyword appears in the text and the word vectors corresponding to each keyword includes:
[0027] Taking the total number of times the keyword appears in the text as the number of rows of the word vector matrix and taking the number of vector dimensions as the number of columns of the word vector matrix, and performing splicing processing on the word vectors corresponding to the keywords respectively to obtain the word vector matrix;
[0028] Among them, the number of times the word vector corresponding to any keyword appears in the word vector matrix is the same as the number of times it appears in the text.
[0029] In an alternative embodiment, fusing the first word vector matrix and the second word vector matrix to obtain a fusion matrix includes:
[0030] Performing matrix multiplication on the first word vector matrix and the second word vector matrix to obtain the fusion matrix.
[0031] In an alternative embodiment, the method further includes:
[0032] When the first text belongs to the text of interest of the target user and the relevance between the second text and the first text corresponding to each of the multiple vector dimensions meets a preset condition, taking the second text as the text of interest of the target user.
[0033] In a second aspect, an embodiment of the present disclosure further provides a text processing device, including:
[0034] An acquisition module, configured to acquire a first word vector matrix corresponding to a first keyword in a first text and a second word vector matrix corresponding to a second keyword in a second text; where, the word vector matrix includes: word vectors corresponding to multiple keywords; each word vector includes: vector elements corresponding to multiple vector dimensions;
[0035] A processing module, configured to fuse the first word vector matrix and the second word vector matrix to obtain a fusion matrix, and perform singular value decomposition on the fusion matrix to obtain a singular value matrix;
[0036] A determination module, configured to determine the correlation degrees of the first text and the second text in multiple vector dimensions based on the singular value matrix.
[0037] In a third aspect, an embodiment of the present disclosure further provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps in the above first aspect or any possible implementation manner in the first aspect are executed.
[0038] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps in the above first aspect or any possible implementation manner in the first aspect are executed.
[0039] In the embodiment of the present disclosure, the first word vector matrix corresponding to the first keyword in the first text and the second word vector matrix corresponding to the second keyword in the second text are fused to obtain the relevant word vectors of the first text and the second text, that is, the vectors included in the fusion matrix. By performing singular value decomposition on the fusion matrix, a singular value matrix corresponding to each matrix dimension of the compressed word vector matrix (that is, the number of compressed word vectors) can be obtained. Each singular value included in the singular value matrix can reflect the correlation degree between the compressed word vector corresponding to the first text and the compressed word vector corresponding to the second text in each matrix dimension. Furthermore, the overall correlation between the first text and the second text can be more accurately and comprehensively characterized through the correlation degrees in multiple matrix dimensions.
[0040] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and are described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for the embodiments will be briefly introduced below. The accompanying drawings are incorporated into the specification and form a part of this specification. These accompanying drawings show embodiments consistent with the present disclosure and are used together with the specification to explain the technical solutions of the present disclosure. It should be understood that the following accompanying drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related accompanying drawings can be obtained based on these accompanying drawings without creative efforts.
[0042] Figure 1 Shows a flowchart of a text processing method provided by an embodiment of the present disclosure;
[0043] Figure 2 shows a flowchart of constructing a word vector matrix provided by an embodiment of the present disclosure;
[0044] Figure 3 shows a flowchart of another text processing method provided by an embodiment of the present disclosure;
[0045] Figure 4 shows a schematic diagram of a text processing device provided by an embodiment of the present disclosure;
[0046] Figure 5 shows a schematic diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. Components of the embodiments of the present disclosure described and illustrated herein generally may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0048] In the process of determining the overall relevance of different texts, generally, the words in different texts are respectively transformed into word vectors, and for each text, a word vector matrix corresponding to the text is respectively constructed by using the word vectors in each text. Then, the mean value between different word vectors in each word vector matrix is calculated to obtain the compressed word vectors corresponding to each text. Next, the cosine similarity between the compressed word vectors corresponding to different texts is calculated. The overall similarity between different texts is determined through the cosine similarity between the compressed word vectors. The above method for determining the relevance of different texts has the problem that the relevance of different texts cannot be accurately reflected.
[0049] Based on this, the present disclosure provides a text processing method, including: obtaining a first word vector matrix corresponding to a first keyword in a first text and a second word vector matrix corresponding to a second keyword in a second text; performing a fusion process on the first word vector matrix and the second word vector matrix to obtain a fusion matrix, and performing singular value decomposition on the fusion matrix to obtain a singular value matrix; determining the relevance between the first text and the second text in multiple vector dimensions based on the singular value matrix. In the above process, by performing a fusion process on the first word vector matrix corresponding to the first keyword in the first text and the second word vector matrix corresponding to the second keyword in the second text, a relevant word vector of the first text and the second text is obtained, that is, the vector included in the fusion matrix; by performing singular value decomposition on the fusion matrix, a singular value matrix corresponding to each matrix dimension of the compressed word vector matrix (i.e., the number of compressed word vectors) can be obtained, and each singular value included in the singular value matrix can reflect the relevance between the compressed word vector corresponding to the first text and the compressed word vector corresponding to the second text in each matrix dimension. Furthermore, the overall relevance between the first text and the second text can be characterized more accurately and comprehensively through the relevance in multiple matrix dimensions.
[0050] Regarding the defects existing in the above solutions and the proposed solutions, they are all the results obtained by the inventors through practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the present disclosure for the above problems in the following text should all be the contributions made by the inventors to the present disclosure during the process of the present disclosure.
[0051] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0052] To facilitate the understanding of this embodiment, first, a text processing method disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the text processing method provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities.
[0053] The text processing method provided in the embodiments of the present disclosure will be described below.
[0054] See Figure 1 As shown, it is a flowchart of a text processing method provided in an embodiment of the present disclosure. The method includes S101 to S103, where:
[0055] S101: Obtain a first word vector matrix corresponding to a first keyword in a first text and a second word vector matrix corresponding to a second keyword in a second text; wherein, the word vector matrix includes: word vectors respectively corresponding to multiple keywords; each of the word vectors includes: vector elements respectively corresponding to multiple vector dimensions.
[0056] In the embodiments of the present disclosure, the first text and the second text may refer to target texts for determining the relevance between texts. The first text and the second text may be, for example, two articles, or may be two groups of comments under the same article. The keywords included in the text may refer to keywords with semantics. The word vector matrix may be constructed based on the word vectors of the keywords. Each keyword's word vector may correspond to multiple vector dimensions, for example, it may be 256-dimensional or 128-dimensional.
[0057] The process of obtaining the word vector matrix corresponding to the keyword in the text, as Figure 2 shown, may specifically include: obtaining the text, and then obtaining the keywords in the text; determining the word vectors corresponding to the keywords; and constructing the word vector matrix corresponding to the text based on the word vectors corresponding to the keywords.
[0058] In practice, the keywords in the text can be obtained through a preprocessing process. The preprocessing process will be described in detail below.
[0059] In the process of determining the word vectors corresponding to the keywords, in one implementation, the word vectors corresponding to the keywords can be determined according to the pre-generated correspondence between the keywords and the word vectors. Here, a word vector library can be maintained, and the correspondence between the keywords and the word vectors can be stored in the word vector library. The keywords here can be the words in a preset vocabulary library. In another implementation, a pre-trained natural language processing model can be used to convert the keywords in the input text into corresponding word vectors.
[0060] For the first text, the word vectors corresponding to the first keyword in the first text can be concatenated to construct a first word vector matrix corresponding to the first text; for the second text, the word vectors corresponding to the second keyword in the second text can be concatenated to construct a second word vector matrix corresponding to the second text.
[0061] Continuing from the previous text, the preprocessing process of the text may include word segmentation processing, keyword screening processing, etc. By preprocessing the text, the influence of the noise contained in the text on the text processing result can be reduced. In one implementation, in the process of obtaining the word vector matrix corresponding to the keyword in the text, the text can be preprocessed as described above first, and then the word vector matrix corresponding to the text can be constructed based on the keywords obtained after the preprocessing.
[0062] Specifically, the text can be tokenized first to obtain the initial words contained in the text; then, the initial words can be filtered according to the preset filtering rules to obtain the keywords; next, a word vector matrix corresponding to the text can be constructed based on the number of times each keyword appears in the text and the word vectors corresponding to each keyword.
[0063] Tokenizing the text means splitting the text into individual initial words. The initial words can include words with semantics and words without semantics. Words without semantics include, for example, auxiliary words and modal particles.
[0064] The preset filtering rules can include filtering words with semantics. By filtering words with semantics, each keyword in the text can be obtained. This process can also be called stop-word removal, that is, removing the words without semantics in the text.
[0065] The preset filtering rules can also include filtering words that appear at least once. That is, during the filtering process, the keywords can be not de-duplicated. In the case of not de-duplicating, the number of the above keywords can be equal to the total number of times the keywords appear. The reason for not de-duplicating the keywords here is mainly that the more times a certain keyword appears in the text, the closer the keyword is to the theme information of the text, that is, the more important the keyword is. For example, in the case of not de-duplicating, the total number of keywords contained in a certain text is 1,000, and the keyword "education" appears 800 times, that is, "education" appears 800 times in this text, indicating that the proportion of this keyword in this text is relatively high, so it can be considered that this keyword is a relatively important keyword.
[0066] In order to further reduce text noise, in one implementation, the spaces, symbols and other noises in the text can also be removed before tokenizing the text.
[0067] For example, the text before preprocessing includes: "74.6% of the students, especially those with medium or average grades, sighed that the exam time was short, there were too many questions to finish, there were many pitfalls in the questions, and it was easy to make mistakes. Some students also reported that they could understand the teacher's explanation, but got stuck when doing the questions and didn't know how to start."; The text obtained after removing noises such as punctuation marks can be: "of the students especially those with medium or average grades all sighed that the exam time was short there were too many questions to finish there were many pitfalls in the questions it was easy to make mistakes some students also reported that they could understand the teacher's explanation but got stuck when doing the questions didn't know how to start"; After word segmentation, the initial words obtained can be: " / of / the / students / especially / grades / medium / or / average / of / the / students / all / sighed / the / exam / time / short / questions / do / not / finish / questions / pitfalls / many / very / easy / make / mistakes / also / some / students / reported / listen / teacher / speak / can / listen / understand / but / one / do / questions / then / get / stuck / do / not / know / how / start"; The keywords obtained after screening can be: "students / especially / grades / medium / average / students / sighed / the / exam / time / short / questions / questions / pitfalls / many / easy / make / mistakes / students / reported / listen / teacher / speak / listen / understand / do / questions / get / stuck / do / not / know / how / start".
[0068] In specific implementation, the above preprocessing process can be respectively executed for the first text and the second text to obtain the first keywords in the first text and the second keywords in the second text. Next, a first word vector matrix corresponding to the first keywords in the first text and a second word vector matrix corresponding to the second keywords in the second text can be constructed.
[0069] In the process of constructing the word vector matrix, the order of appearance of the keywords in the text may not be considered, and only the number of times the keywords appear in the text needs to be obtained. Of course, in other implementation manners, the order of appearance of the keywords in the text can also be considered, that is, a word vector matrix corresponding to the text can be constructed based on the order of appearance of each keyword in the text, the number of times of appearance, and the word vectors corresponding to each keyword. In this case, the order of each word vector in the constructed word vector matrix can be consistent with the order of the keywords in the text.
[0070] When constructing an initial word vector matrix corresponding to the text based on the number of times each keyword appears in the text respectively and the word vectors corresponding to each keyword, in one implementation manner, the following steps may be included: taking the total number of times the keyword appears in the text as the number of rows of the word vector matrix and taking the number of vector dimensions as the number of columns of the word vector matrix, and performing splicing processing on the word vectors corresponding to the keywords respectively to obtain the word vector matrix; wherein, the number of times the word vector corresponding to any keyword appears in the word vector matrix is the same as the number of times it appears in the text.
[0071] For example, in the first text, the total number of occurrences of each first keyword is n1 times, and the number of vector dimensions of the word vectors corresponding to each first keyword is P; in the second text, the total number of occurrences of each second keyword is n2 times, and the number of vector dimensions of the word vectors corresponding to each second keyword is also P.
[0072] For the first text, n1 can be used as the number of rows of the word vector matrix, and P can be used as the number of columns of the word vector matrix. The word vectors corresponding to the first keywords are concatenated respectively to obtain a first word vector matrix with n1 rows and P columns. For the second text, n2 can be used as the number of rows of the word vector matrix, and P can be used as the number of columns of the word vector matrix. The word vectors corresponding to the second keywords are concatenated respectively to obtain a second word vector matrix with n2 rows and P columns.
[0073] In other embodiments, the total number of occurrences of the keywords in the text can also be used as the number of columns of the word vector matrix, and the number of vector dimensions can be used as the number of rows of the word vector matrix. The word vectors corresponding to the keywords are concatenated respectively to obtain a word vector matrix.
[0074] In this embodiment, for the first text in the above example, a first word vector matrix with P rows and n1 columns can be obtained; for the second text in the above example, a second word vector matrix with P rows and n2 columns can be obtained.
[0075] S102: Perform a fusion process on the first word vector matrix and the second word vector matrix to obtain a fusion matrix, and perform a singular value decomposition on the fusion matrix to obtain a singular value matrix.
[0076] In the embodiments of the present disclosure, this step can be a canonical correlation process on the first word vector matrix and the second word vector matrix. First, a fusion process can be performed on the first word vector matrix and the second word vector matrix to obtain a fusion matrix. In one embodiment, the first word vector matrix can be multiplied by the second word vector matrix, that is, a matrix multiplication process is performed on the first word vector matrix and the second word vector matrix. Then, a singular value decomposition is performed on the fusion matrix.
[0077] In one embodiment, the obtained fusion matrix can have the total number of occurrences of the first keywords in the first text as the number of rows of the fusion matrix, and the total number of occurrences of the second keywords in the second text as the number of columns of the fusion matrix.
[0078] If the word vector matrix is constructed in such a way that the number of total occurrences of a keyword in the text is used as the number of rows of the word vector matrix, and the number of vector dimensions is used as the number of columns of the word vector matrix, then at this time, since the total number of occurrences of the first keyword in the first text may be different from the total number of occurrences of the second keyword in the second text, the first word vector matrix and the second word vector matrix can be transposed. After transposing once, both the first word vector matrix and the second word vector matrix after transposing once use the number of vector dimensions as the number of rows (the number of vector dimensions of the word vectors in the first word vector matrix is the same as the number of vector dimensions of the word vectors in the second word vector matrix), and at this time, it can be ensured that the number of rows of the first word vector matrix and the second word vector matrix is the same.
[0079] Then, transpose the first word vector matrix one more time, and multiply the first word vector matrix after transposing twice by the first word vector matrix after transposing once to obtain the fusion matrix.
[0080] Suppose there is a first word vector matrix X1 with P rows and 1000 columns and a second word vector matrix X2 with P rows and 2000 columns. Among them, the first word vector matrix X1 can represent that the total number of occurrences of the first keyword in the first text is 1000 times, and the vector dimension of the word vector corresponding to the first keyword is P dimensions; the second word vector matrix X2 can represent that the total number of occurrences of the second keyword in the second text is 2000 times, and the vector dimension of the word vector corresponding to the second keyword is also P dimensions.
[0081] Here, the fusion matrix S can be obtained through Specifically, first transpose the first word vector matrix X1 with P rows and 1000 columns to obtain a matrix with 1000 rows and P columns Then multiply the matrix with 1000 rows and P columns by the second word vector matrix X2 with P rows and 2000 columns to obtain the fusion matrix S with 1000 rows and 2000 columns.
[0082] In the above process, the first word vector matrix needs to be transposed twice, and the second word vector matrix needs to be transposed once. Therefore, in one implementation, only the second word vector matrix can be transposed once, and the first word vector matrix can be multiplied by the transposed second word vector matrix to obtain the fusion matrix.
[0083] If the word vector matrix is constructed in such a way that the number of vector dimensions is used as the number of rows of the word vector matrix, and the number of total occurrences of a keyword in the text is used as the number of columns of the word vector matrix, then the first word vector matrix can be transposed once, and the transposed first word vector matrix can be multiplied by the second word vector matrix to obtain the fusion matrix.
[0084] Next, perform singular value decomposition on the fusion matrix to obtain a singular value matrix containing each singular value. After decomposing the fusion matrix, two orthogonal matrices can also be obtained, namely the first target decomposition matrix and the second target decomposition matrix described later. Based on the first target decomposition matrix and the singular value matrix, the first word vector matrix can be compressed to obtain a compressed word vector matrix. Based on the second target decomposition matrix and the singular value matrix, the second word vector matrix can be compressed to obtain another compressed word vector matrix. The matrix dimensions of the two obtained compressed word vector matrices are the same. Here, the matrix dimension is the number of word vectors in the compressed word vector matrix. Each singular value is the element value on the diagonal of the singular value matrix, and each singular value can represent the similarity information between the first semantic keywords corresponding to the word vectors of the first text under each matrix dimension of the compressed word vector matrix and the second semantic keywords corresponding to the word vectors of the second text under each matrix dimension of the compressed word vector matrix.
[0085] S103: Determine the relevance of the first text and the second text respectively in multiple vector dimensions based on the singular value matrix.
[0086] In a specific implementation, the similarity information between the first semantic keywords corresponding to the word vectors of the first text under each matrix dimension corresponding to each singular value in the singular value matrix and the second semantic keywords corresponding to the word vectors of the second text under each matrix dimension can be determined as the relevance of the first text and the second text respectively in multiple vector dimensions.
[0087] In the embodiments of the present disclosure, when performing singular value decomposition on the fusion matrix, the first target decomposition matrix and the second target decomposition matrix can also be obtained. Among them, the first target decomposition matrix is used to represent the weights corresponding to the semantics of the first text in multiple vector dimensions respectively; the second target decomposition matrix is used to represent the weights corresponding to the semantics of the second text in multiple vector dimensions respectively.
[0088] The number of decomposition vectors included in the first target decomposition matrix can be the same as the number of rows of the fusion matrix. The number of decomposition vectors included in the second target decomposition matrix can be the same as the number of columns of the fusion matrix.
[0089] In a specific implementation process, the first semantic keywords of the first text in multiple matrix dimensions can be determined based on the first word vector matrix and the first target decomposition matrix, and the second semantic keywords of the second text in multiple matrix dimensions can be determined based on the first word vector matrix and the second target decomposition matrix.
[0090] Here, since the first target decomposition matrix represents the weights corresponding to the semantics of the first text in multiple vector dimensions, in one implementation, the word vector compression matrix corresponding to the first word vector matrix can be determined based on the first target decomposition matrix and the first word vector matrix. Then, based on each word vector in the word vector compression matrix and the word vectors corresponding to multiple candidate words, the first semantic keywords corresponding to each word vector can be determined. Based on the same process, the second semantic keywords can also be obtained, which will not be elaborated here.
[0091] In the process of determining the word vector compression matrix corresponding to the first word vector matrix based on the first target decomposition matrix and the first word vector matrix, in one implementation, the first target decomposition matrix can be compressed according to a preset number of singular values to obtain the compressed decomposition matrix corresponding to the first target decomposition matrix. Then, based on the compressed decomposition matrix and the first word vector matrix, the word vector compression matrix corresponding to the first word vector matrix can be determined.
[0092] The preset number of singular values can be the first preset number of singular values arranged in the singular value matrix. The maximum value of the preset number can be the minimum value among the vector dimension, the number of keywords in the first text, and the number of keywords in the second text. Here, for each singular value among the first preset number of singular values arranged in the singular value matrix, the decomposition vector with the same number of columns in the first target decomposition matrix as the number of rows (or columns; when the singular value matrix is a diagonal matrix, the number of rows of each singular value on the diagonal is the same as the number of columns) of this singular value can be determined, that is, through the first preset number of singular values arranged in the singular value matrix, the decomposition vectors of the first preset number of columns in the first target decomposition matrix can be taken, so that the decomposition vectors of the first preset number of columns are determined as the compressed decomposition matrix corresponding to the first target decomposition matrix, realizing the compression of the first target decomposition matrix. By multiplying the compressed decomposition matrix and the first word vector matrix, the word vector compression matrix corresponding to the first word vector matrix can be determined.
[0093] Based on a similar process, the word vector compression matrix corresponding to the second word vector matrix can also be obtained, which will not be elaborated here.
[0094] Finally, for each word vector in the word vector compression matrix corresponding to the first word vector matrix, the first semantic keyword corresponding to this word vector is determined from the word vectors corresponding to multiple candidate words, and for each word vector in the word vector compression matrix corresponding to the second word vector matrix, the second semantic keyword corresponding to this word vector is determined from the word vectors corresponding to multiple candidate words.
[0095] The candidate words can be words in a preset vocabulary. The word vectors corresponding to the candidate words can be obtained based on the aforementioned pre-trained natural language processing model.
[0096] Exemplarily, continuing from the previous text, for the first word vector matrix X1 with P rows and 1000 columns and the second word vector matrix X2 with P rows and 2000 columns, through the formula After obtaining the fusion matrix S with 1000 rows and 2000 columns, the fusion matrix S with 1000 rows and 2000 columns can be subjected to singular value decomposition to obtain an orthogonal matrix W1 with 1000 rows and 1000 columns, a singular value matrix V with 1000 rows and 1000 columns, and an orthogonal matrix W2 with 2000 rows and 2000 columns.
[0097] For example, according to the first 3 singular values in the singular value matrix V (here, the number of singular values can be determined when the ratio of the sum of the selected singular value values to the sum of all singular values in the singular value matrix meets the set threshold, and the magnitudes of the singular values in the singular value matrix are arranged in the order from front to back, that is, the more forward singular value has a larger value), the first compression decomposition matrix Y1 with 1000 rows and 3 columns corresponding to the orthogonal matrix W1 can be obtained, and the second compression decomposition matrix Y2 with 2000 rows and 3 columns corresponding to the orthogonal matrix W2 can be obtained.
[0098] Then, multiplying the first word vector matrix X1 with P rows and 1000 columns by the first compression matrix Y1 with 1000 rows and 3 columns, the first word vector compression matrix with P rows and 3 columns corresponding to the first word vector matrix X1 can be obtained; multiplying the second word vector matrix X2 with P rows and 2000 columns by the second compression matrix Y2 with 2000 rows and 3 columns, the second word vector compression matrix with P rows and 3 columns corresponding to the first word vector matrix X2 can be obtained.
[0099] Finally, for each column of word vectors in the first word vector compression matrix, find the first semantic keyword corresponding to the word vector; for each column of word vectors in the second word vector compression matrix, find the second semantic keyword corresponding to the word vector.
[0100] Among them, the singular value corresponding to each column of word vectors can be used as the relevance between the first semantic keyword corresponding to the first text in this column of word vectors and the second semantic keyword corresponding to the second text in this column of word vectors.
[0101] In one implementation, after obtaining the relevance between the first text and the second text in multiple vector dimensions, when the first text belongs to the text of interest of the target user and the relevance between the second text and the first text corresponding to each of the multiple vector dimensions meets the preset conditions, the second text can be used as the text of interest of the target user.
[0102] Here, the text of interest to the target user can be preselected as the first text, or it can be determined that the first text is the text of interest to the target user according to the tags preset for the first text. In either case, when it is determined that the first text belongs to the text of interest to the target user and the relevance of the second text corresponding to the first text in multiple vector dimensions meets the preset conditions, the second text can be used as the text of interest to the target user.
[0103] In one implementation, the relevance thresholds corresponding to the second text and the first text in multiple vector dimensions can be set. The relevance here can be similarity information. In this implementation, the second text with a relevance greater than the relevance threshold corresponding in multiple vector dimensions can be used as the text of interest to the target user. Using the first text and the second text that the target user is interested in, the interest characteristics of the target user can be analyzed, and the above-mentioned second text, or texts of the same type as the first text and the second text, etc., can be pushed to the target user. The specific applications of the first text and the second text here can be not specifically limited.
[0104] The embodiments of the present disclosure also provide a flowchart of another text processing method, as Figure 3 shown.
[0105] And Figure 3 the corresponding specific process, for example, includes the following steps a1 to step a4:
[0106] Step a1: The first original text and the second original text can be obtained, for example, two sets of review texts of an article.
[0107] Step a2: The preprocessing module is used to preprocess the first original text and the second original text respectively (here, it can be word segmentation and stop word removal, etc.), and the first keyword in the first original text and the second keyword in the second original text can be obtained; next, using the corresponding relationship between the preset keywords and word vectors provided by the pre-trained model, the first keyword and the second keyword are respectively transformed into word vectors, and the word vectors corresponding to the first keyword are spliced into a first word vector matrix, and the word vectors corresponding to the second keyword are spliced into a second word vector matrix. Among them, the process of splicing word vectors into a word vector matrix can refer to the previous text and will not be elaborated here.
[0108] Step a3: Perform canonical correlation processing on the first word vector matrix and the second word vector matrix.
[0109] Specifically, the first word vector matrix can be multiplied by the second word vector matrix to obtain a fusion matrix. Next, singular value decomposition is performed on the fusion matrix to obtain a first target decomposition matrix for characterizing the weight coefficients corresponding to the semantics of the first original text in multiple vector dimensions, a second target decomposition matrix for characterizing the weight coefficients corresponding to the semantics of the second original text in multiple vector dimensions, and a singular value matrix; wherein, the singular value matrix contains multiple singular values, and the singular value is the value of the matrix element corresponding to the target vector dimension. Here, the process of performing singular value decomposition on the fusion matrix can refer to the previous text and will not be elaborated here. Next, according to the first preset number of singular values in the singular value matrix, the first target decomposition matrix is compressed to obtain a first compressed decomposition matrix, and the first compressed decomposition matrix is multiplied by the first word vector matrix to obtain a first word vector compression matrix corresponding to the first word vector matrix; according to the first preset number of singular values in the singular value matrix, the second target decomposition matrix is compressed to obtain a second compressed decomposition matrix, and the second compressed decomposition matrix is multiplied by the second word vector matrix to obtain a second word vector compression matrix corresponding to the second word vector matrix.
[0110] Step a4: Determine the relevance between the first original text and the second original text, the first semantic keyword, and the second semantic keyword.
[0111] The first preset number of singular values in the singular value matrix represent the relevance between the first original text and the second original text in each vector dimension.
[0112] Using the correspondence between the candidate keywords provided by the pre-trained model and the word vectors, and the word vectors corresponding to each semantics in the first word vector compression matrix, the first semantic keywords whose similarity information with each semantics meets the set threshold are screened out from the candidate keywords. Using the correspondence between the candidate keywords provided by the pre-trained model and the word vectors, and the word vectors corresponding to each semantics in the second word vector compression matrix, the second semantic keywords whose similarity information with each semantics meets the set threshold are screened out from the candidate keywords.
[0113] The following is the process of processing the text based on the text processing provided in the embodiments of the present disclosure. Specifically, the first text obtained, for example, the first group of review texts of an article is as follows:
[0114] "There are many. Among more than four million people taking the postgraduate entrance examination, only more than one million are admitted. There is no way. If you decide to take the postgraduate entrance examination, everyone is very hardworking, but the total number of admissions is there.
[0115] Thinking of my own preparation process, from the sweltering heat to the cold winter, how many sunrises and sunsets, just hoping for a result. I can only say that I have a result this year. If I fail, I have no way out, I have no choice, and I am not calm. Everyone is advising, advising that there is a way ahead, but how do you know there is a way ahead if you haven't really experienced it.
[0116] [Crying uncontrollably] Seeing the hardships these people have gone through in the postgraduate entrance examination, I feel a resonance, but more of a numbness. [Tears streaming down my face]
[0117] Moreover, everything is becoming more and more competitive. The cut-off scores for this year's postgraduate entrance examination have even increased.
[0118] Standing at this point, it's so difficult to take the postgraduate entrance examination, the civil service examination, the establishment examination, or to find a job. [Crying][Crying][Crying]
[0119] The second text obtained, for example, the second set of review texts of an article is as follows:
[0120] "Come on! Everyone who perseveres in walking this postgraduate entrance examination path deserves to applaud themselves. When the length of life is long enough, the result at a certain moment is not that important. Believe that the path you have walked is not in vain. Bless you to meet a prosperous future at another intersection. [Love]"
[0121] I took the exam for three years. When the results came out, I felt like a joke. How many all-nighters did I stay up, how many notebooks and mind maps did I organize, how many questions did I brush, and how many books did I memorize. Really, when the results came out, I felt like a big joke. I'm already thirty years old. Really, it's time to give up.
[0122] I took the 311 exam. Both my politics and English scores exceeded 70, but my professional course scores were about the same as when I only prepared for two months in the first year. It doesn't matter. Those who work hard won't have bad luck! I took the Tianjin Normal University exam in 2020. I recited the professional course 6 times but still didn't reach 100 points, but I got into the interview. Later, I still didn't get in! However, in the same year, I took the teacher establishment exam as a fresh graduate! Later, during the process of studying for the teacher establishment exam, a teacher who taught pedagogy made me figure it out too.
[0123] This requires 6 books to be interconnected. If you plan to take the exam for the second time, remember to integrate the content of the 6 books to answer the big questions! Using this method, I passed the teacher establishment exam twice and chose one of them! Keep working hard.
[0124] After performing word segmentation and stop word removal on the above first set of review texts, the following keywords can be obtained: "postgraduate entrance examination / total / one million / helpless / postgraduate entrance examination / hard work / enrollment / total number / there / preparation process / sweltering heat / winter / result / failure / way out / choice / calm / experience / postgraduate entrance examination / hardship / resonance / numbness / intense competition / cut-off score / increase / civil service examination / establishment examination / find a job / difficult".
[0125] After performing word segmentation and stop word removal on the above second set of review texts, the following keywords can be obtained: "come on / persevere / walk through / postgraduate entrance examination path / applaud / life / result / in vain / prosperous future / score / joke / all-nighter / note / give up / professional course / luck / fresh graduate / identity / teacher establishment exam / teach / pedagogy / teacher / take the exam for the second time / content / integrate / teacher establishment exam".
[0126] Next, a pre-trained model can be used to convert the keywords in the first group of review texts and the keywords in the second group of review texts into word vectors respectively. Then, the word vectors corresponding to the keywords in the first group of review texts are concatenated into a first word vector matrix; the word vectors corresponding to the keywords in the second group of review texts are concatenated into a second word vector matrix.
[0127] Next, the above-mentioned canonical correlation processing is performed on the first word vector matrix and the second word vector matrix, and the following results are obtained: the similarity between the first group of review texts and the second group of review texts in the compressed first vector dimension is 0.9939; the information after compression of the first group of review texts is: "Education / Family Education / Education Science"; the information after compression of the second group of review texts is: "Education / Family Education / Education Science".
[0128] The similarity between the first group of review texts and the second group of review texts in the compressed second vector dimension is 0.9957; the information after compression of the first group of review texts is: "Postgraduate Entrance Examination / Preparing for the Exam / Classmates"; the information after compression of the second group of review texts is: "Postgraduate Entrance Examination / Preparing for the Exam / Classmates".
[0129] The similarity between the first group of review texts and the second group of review texts in the compressed third vector dimension is 0.7677; the information after compression of the first group of review texts is: "Grand View Garden / Colored Sand / Flight Altitude"; the information after compression of the second group of review texts is: "Education / Normal Education / Special Education".
[0130] It can be seen that the semantic keywords obtained above describe the relevance between the first group of review texts and the second group of review texts from different perspectives.
[0131] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0132] Based on the same inventive concept, a text processing device corresponding to the text processing method is further provided in the embodiments of the present disclosure. Since the principle of solving problems by the device in the embodiments of the present disclosure is similar to the above-mentioned text processing method in the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0133] Refer to Figure 4 As shown, it is a schematic architecture diagram of a text processing device provided by an embodiment of the present disclosure. The device includes: an acquisition module 401, a processing module 402, and a determination module 403; wherein,
[0134] An acquisition module 401 is configured to acquire a first word vector matrix corresponding to a first keyword in a first text and a second word vector matrix corresponding to a second keyword in a second text; wherein, the word vector matrix includes: word vectors respectively corresponding to a plurality of keywords; each of the word vectors includes: vector elements respectively corresponding to a plurality of vector dimensions;
[0135] A processing module 402 is configured to perform a fusion process on the first word vector matrix and the second word vector matrix to obtain a fusion matrix, and perform a singular value decomposition on the fusion matrix to obtain a singular value matrix;
[0136] A determination module 403 is configured to determine the relevance of the first text and the second text respectively in a plurality of the vector dimensions based on the singular value matrix.
[0137] In an optional implementation manner, when performing the singular value decomposition on the fusion matrix, a first target decomposition matrix and a second target decomposition matrix are further obtained; the first target decomposition matrix is used to represent the weights respectively corresponding to the semantics of the first text in a plurality of the vector dimensions; the second target decomposition matrix is used to represent the weights respectively corresponding to the semantics of the second text in a plurality of the vector dimensions;
[0138] The determination module 403 is specifically configured to determine first semantic keywords of the first text respectively in a plurality of the vector dimensions based on the first word vector matrix and the first target decomposition matrix, and determine second semantic keywords of the second text respectively in a plurality of the vector dimensions based on the first word vector matrix and the second target decomposition matrix.
[0139] In an optional implementation manner, the determination module 403 is specifically configured to:
[0140] Determine a word vector compression matrix corresponding to the first word vector matrix based on the first word vector matrix and the first target decomposition matrix;
[0141] Determine first semantic keywords corresponding to each of the word vectors based on each of the word vectors in the word vector compression matrix and word vectors respectively corresponding to a plurality of candidate words.
[0142] In an optional implementation manner, the determination module 403 is specifically configured to:
[0143] Compress the first target decomposition matrix according to a preset number of singular values to obtain a compressed decomposition matrix corresponding to the first target decomposition matrix; wherein, the singular value is the value of the matrix element corresponding to the target vector dimension in the singular value matrix;
[0144] Determine the word vector compression matrix corresponding to the first word vector matrix based on the compression decomposition matrix and the first word vector matrix.
[0145] In an optional implementation, the obtaining module 401 is specifically configured to:
[0146] Perform word segmentation on the text to obtain the initial words included in the text;
[0147] Perform screening processing on the initial words according to a preset screening rule to obtain the keyword;
[0148] Construct a word vector matrix corresponding to the text based on the number of times each keyword appears in the text and the word vectors corresponding to each keyword;
[0149] Wherein, the text includes a first text, the keyword includes a first keyword, and the word vector matrix includes the first word vector matrix;
[0150] Or,
[0151] The text includes a second text, the keyword includes a second keyword, and the word vector matrix includes the second word vector matrix.
[0152] In an optional implementation, the obtaining module 401 is specifically configured to:
[0153] Use the total number of times the keyword appears in the text as the number of rows of the word vector matrix, and use the number of vector dimensions as the number of columns of the word vector matrix, and splice the word vectors corresponding to the keywords respectively to obtain the word vector matrix;
[0154] Wherein, the number of times the word vector corresponding to any keyword appears in the word vector matrix is the same as the number of times it appears in the text.
[0155] In an optional implementation, the processing module 402 is specifically configured to:
[0156] Perform matrix multiplication on the first word vector matrix and the second word vector matrix to obtain the fusion matrix.
[0157] In an optional implementation, the device further includes:
[0158] A second determination module, configured to use the second text as the text of interest of the target user when the first text belongs to the text of interest of the target user and the relevance of the second text corresponding to the first text in multiple vector dimensions meets a preset condition.
[0159] Descriptions of the processing flows of the various modules in the device and the interaction flows between the various modules can refer to the relevant descriptions in the foregoing method embodiments and will not be elaborated here.
[0160] Based on the same inventive concept, embodiments of the present disclosure also provide a computer device. Referring to Figure 5 As shown, it is a schematic structural diagram of a computer device 500 provided by an embodiment of the present disclosure, including a processor 501, a memory 502, and a bus 503. Among them, the memory 502 is used to store execution instructions, including an internal memory 5021 and an external memory 5022; the internal memory 5021 here is also called the main memory, which is used to temporarily store the operation data in the processor 501 and the data exchanged with the external memory 5022 such as a hard disk. The processor 501 exchanges data with the external memory 5022 through the internal memory 5021. When the computer device 500 runs, the processor 501 communicates with the memory 502 through the bus 503, so that the processor 501 executes the following instructions:
[0161] Obtain a first word vector matrix corresponding to a first keyword in a first text and a second word vector matrix corresponding to a second keyword in a second text; wherein, the word vector matrix includes: word vectors respectively corresponding to multiple keywords; each of the word vectors includes: vector elements respectively corresponding to multiple vector dimensions;
[0162] Perform a fusion process on the first word vector matrix and the second word vector matrix to obtain a fusion matrix, and perform a singular value decomposition on the fusion matrix to obtain a singular value matrix;
[0163] Determine the relevance of the first text and the second text in multiple vector dimensions based on the singular value matrix.
[0164] Embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the text processing method described in the foregoing method embodiments. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.
[0165] Embodiments of the present disclosure also provide a computer program product. The computer product carries program codes, and the instructions included in the program codes can be used to execute the steps of the text processing method described in the foregoing method embodiments. For details, refer to the foregoing method embodiments and will not be elaborated here.
[0166] Among them, the above computer program product can be specifically implemented by means of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0167] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein. In several embodiments provided in the present disclosure, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0168] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0169] In addition, in each embodiment of the present disclosure, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0170] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0171] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A text processing method, characterized in that, Including: Obtain a first word vector matrix corresponding to a first keyword in a first text and a second word vector matrix corresponding to a second keyword in a second text; wherein, the word vector matrix includes: word vectors respectively corresponding to a plurality of keywords; each of the word vectors includes: vector elements respectively corresponding to a plurality of vector dimensions; Perform a fusion process on the first word vector matrix and the second word vector matrix to obtain a fusion matrix, and perform a singular value decomposition on the fusion matrix to obtain a first target decomposition matrix and a second target decomposition matrix; the first target decomposition matrix is used to represent the weights respectively corresponding to the semantics of the first text under a plurality of the vector dimensions; the second target decomposition matrix is used to represent the weights respectively corresponding to the semantics of the second text under a plurality of the vector dimensions; The method further includes: Based on the first word vector matrix and the first target decomposition matrix, determine first semantic keywords of the first text respectively under a plurality of the vector dimensions, and, based on the second word vector matrix and the second target decomposition matrix, determine second semantic keywords of the second text respectively under a plurality of the vector dimensions; and Determine the similarity information between the first semantic keywords of the first text respectively under a plurality of the vector dimensions and the second semantic keywords of the second text respectively under a plurality of the vector dimensions as the relevance between the first text and the second text respectively under a plurality of the vector dimensions.
2. The method according to claim 1, characterized in that The determining, based on the first word vector matrix and the first target decomposition matrix, the first semantic keywords of the first text respectively under a plurality of the vector dimensions includes: Based on the first word vector matrix and the first target decomposition matrix, determine a word vector compression matrix corresponding to the first word vector matrix; Based on each word vector in the word vector compression matrix and the word vectors respectively corresponding to a plurality of candidate words, determine the first semantic keyword corresponding to each of the word vectors.
3. The method according to claim 2, wherein When performing a singular value decomposition on the fusion matrix, a singular value matrix is further obtained; The determining, based on the first word vector matrix and the first target decomposition matrix, the word vector compression matrix corresponding to the first word vector matrix includes: Compress the first target decomposition matrix according to a preset number of singular values to obtain a compressed decomposition matrix corresponding to the first target decomposition matrix; wherein, the singular value is the value of the matrix element corresponding to a target vector dimension in the singular value matrix; Based on the compressed decomposition matrix and the first word vector matrix, determine the word vector compression matrix corresponding to the first word vector matrix.
4. The method according to claim 1, characterized in that, The obtaining the word vector matrix corresponding to the keyword in the text includes: Perform word segmentation on the text to obtain initial words included in the text; Perform a screening process on the initial words according to a preset screening rule to obtain the keywords; Based on the number of times each of the keywords appears in the text and the word vectors respectively corresponding to each of the keywords, construct the word vector matrix corresponding to the text; Wherein, the text includes a first text, the keyword includes a first keyword, and the word vector matrix includes the first word vector matrix; or, the text includes a second text, the keyword includes a second keyword, and the word vector matrix includes the second word vector matrix.
5. The method according to claim 4, wherein Constructing the word vector matrix corresponding to the text based on the number of times each of the keywords appears in the text and the word vectors corresponding to the keywords respectively includes: Taking the total number of times the keyword appears in the text as the number of rows of the word vector matrix and taking the number of vector dimensions as the number of columns of the word vector matrix, and performing splicing processing on the word vectors corresponding to the keywords respectively to obtain the word vector matrix; Wherein, the number of times the word vector corresponding to any keyword appears in the word vector matrix is the same as the number of times it appears in the text.
6. The method according to claim 1, wherein The fusing the first word vector matrix and the second word vector matrix to obtain a fusion matrix includes: Performing matrix multiplication on the first word vector matrix and the second word vector matrix to obtain the fusion matrix.
7. The method according to any one of claims 1 to 6, characterized in that The method further includes: When the first text belongs to the text of interest of the target user and the relevance degrees of the second text corresponding to the first text in multiple vector dimensions meet preset conditions, taking the second text as the text of interest of the target user.
8. A text processing device, characterized in that, It includes: An acquisition module, configured to acquire a first word vector matrix corresponding to a first keyword in a first text and a second word vector matrix corresponding to a second keyword in a second text; wherein, the word vector matrix includes: word vectors corresponding to multiple keywords respectively; each of the word vectors includes: vector elements corresponding to multiple vector dimensions respectively; A processing module, configured to fuse the first word vector matrix and the second word vector matrix to obtain a fusion matrix, and perform singular value decomposition on the fusion matrix to obtain a first target decomposition matrix and a second target decomposition matrix; the first target decomposition matrix is used to represent the weights corresponding to the semantics of the first text in multiple vector dimensions respectively; the second target decomposition matrix is used to represent the weights corresponding to the semantics of the second text in multiple vector dimensions respectively; Wherein, the apparatus further includes a determination module, configured to: Based on the first word vector matrix and the first target decomposition matrix, determine first semantic keywords of the first text in multiple vector dimensions respectively, and based on the second word vector matrix and the second target decomposition matrix, determine second semantic keywords of the second text in multiple vector dimensions respectively; and Determine the relevance degrees of the first text and the second text in multiple vector dimensions respectively based on the similarity information between the first semantic keywords of the first text in multiple vector dimensions respectively and the second semantic keywords of the second text in multiple vector dimensions respectively.
9. A computer device, characterized in that, It includes: A processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the text processing method according to any one of claims 1 to 7 are performed.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the text processing method according to any one of claims 1 to 7 are performed.
Citation Information
Patent Citations
Retrieval method and method using same for establishing text semantic extraction module
CN102214180A
Online text label real-time adding method and device and related equipment
CN110795911A
A method for summarizing text through sentence extraction
CN110892400A