Word Vector Matrix Enhancement Method, Device, Equipment and Medium
Data enhancement of the word vector matrix through relational transmission solves the problem that word vector matrix is difficult to reflect the semantic information of word vectors in the prior art, and achieves better word vector clustering analysis effect.
Patent Information
- Application Number
- CN202211317581.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-10-26
AI Technical Summary
When constructing a word vector matrix in the prior art, it is difficult to effectively reflect the semantic information of words with low frequency, and the word2vec-based method requires a fixed context window size and cannot adapt to the needs of different scenarios.
By obtaining the text set, the initial word vector matrix is determined, and the matrix elements whose quantized value is the first set value is enhanced based on the relationship transfer method to obtain the target word vector matrix. This method can better reflect the semantic information of the word by enhancing the elements of the word vector matrix.
It realizes effective enhancement of word vectors of words with low occurrence frequency, improves the expressiveness of word vectors in clustering analysis, and is suitable for word vector construction requirements in different scenarios.
Smart Images

Figure CN115599916B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method, device, equipment and medium for enhancing a word vector matrix. Background Art
[0002] With the development of deep learning technology, the performance of natural language processing (NLP) tasks has been greatly improved. Among them, NLP tasks may include: word segmentation, part-of-speech tagging, named entity recognition, sentence classification, sentiment analysis, etc.
[0003] In a document collection composed of many documents, there may be many words with low occurrence frequencies, and the distance (such as Euclidean distance) between the word vectors composed of these words and the word vectors composed of words that do not co-occur in the same document is generally short, and it is difficult to reflect the true semantic information of the words in clustering analysis. Therefore, a method for enhancing word vector data is needed to address this situation.
[0004] Existing methods usually build a word vector matrix based on the neural network method of word2vec. However, the neural network method based on word2vec needs to fix the context window size to train word vectors, which is not applicable when the context window does not need to be fixed, and the word vectors obtained by training the neural network model based on word2vec cannot well explain the features of each dimension in the word vectors. Summary of the Invention
[0005] The present invention provides a method, device, equipment and medium for enhancing text data, which performs data enhancement on the word vector matrix in a relationship transmission manner, so that the word vectors composed of words with low occurrence frequencies can better perform clustering analysis.
[0006] According to one aspect of the present invention, there is provided a method for enhancing a word vector matrix, including:
[0007] Obtain a text set;
[0008] Determine an initial word vector matrix corresponding to the text set; wherein, the word vector matrix is a matrix of M*N, M is the number of word vectors, N is the dimension of the word vectors, and the dimension of the word vectors is equal to the number of texts;
[0009] Perform enhancement processing on matrix elements with a quantization value of a first set value based on the initial word vector matrix to obtain a target word vector matrix.
[0010] Optionally, determining the initial word vector matrix corresponding to the text set includes:
[0011] Obtain a word to be processed;
[0012] For the i-th word, if the i-th word does not appear in the j-th text, the quantization value of the element in the i-th row and j-th column of the initial word vector matrix is determined as the first set value;
[0013] If the i-th word appears in the j-th text, the quantization value of the element in the i-th row and j-th column of the initial word vector matrix is determined as the second set value.
[0014] Optionally, enhancing the matrix elements with the quantization value of the first set value based on the initial word vector matrix includes:
[0015] If the quantization value of the element in the a-th row and b-th column of the initial word vector matrix is the first set value, obtain N quantization values in the a-th word vector as the first quantization values;
[0016] Obtain the quantization values of M words corresponding to the b-th text as the second quantization values;
[0017] Enhance the element in the a-th row and b-th column based on the first quantization values, the second quantization values, and the initial word vector matrix.
[0018] Optionally, enhancing the element in the a-th row and b-th column based on the first quantization values, the second quantization values, and the initial word vector matrix includes:
[0019] Multiply the N first quantization values by the second set value and sum them to obtain a sum value;
[0020] Perform dot products of the a-th word vector with each word vector in the initial word vector matrix to obtain M dot product results;
[0021] Perform weighted summation of the M dot product results based on the M second quantization values;
[0022] Divide the weighted summation result by the sum value to obtain the enhanced quantization value of the element in the a-th row and b-th column.
[0023] Optionally, enhancing the matrix elements with the quantization value of the first set value based on the initial word vector matrix is calculated according to the following formula:
[0024]
[0025] Among them, M represents the number of word vector data to be processed, N represents the number of texts in the text set, H a represents the word vector of the a-th word, and x ab represents the value at the b-th position in the word vector data of the a-th word; A represents the second set value.
[0026] According to another aspect of the present invention, there is provided a word vector matrix enhancement device, including:
[0027] A text set acquisition module for acquiring a text set;
[0028] An initial word vector matrix determination module for determining an initial word vector matrix corresponding to the text set; wherein, the word vector matrix is a matrix of M*N, M is the number of word vectors, N is the dimension of the word vectors, and the dimension of the word vectors is equal to the number of texts;
[0029] A target word vector matrix obtaining module for enhancing matrix elements with a quantization value of a first set value based on the initial word vector matrix to obtain a target word vector matrix.
[0030] Optionally, the initial word vector matrix determination module includes:
[0031] A to-be-processed word acquisition unit for acquiring a to-be-processed word;
[0032] A first set value determination unit for, for the i-th word, if the i-th word does not appear in the j-th text, determining the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as the first set value;
[0033] A second set value determination unit for, if the i-th word appears in the j-th text, determining the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as the second set value.
[0034] Optionally, the target word vector matrix obtaining module includes:
[0035] A first quantization value acquisition unit for, if the quantization value of the element in the a-th row and b-th column of the initial word vector matrix is the first set value, acquiring N quantization values in the a-th word vector as the first quantization value;
[0036] A second quantization value acquisition unit for acquiring quantization values of M words corresponding to the b-th text as the second quantization value;
[0037] An element enhancement processing unit for enhancing the element in the a-th row and b-th column based on the first quantization value, the second quantization value, and the initial word vector matrix.
[0038] According to another aspect of the present invention, there is provided an electronic device, the electronic device including:
[0039] At least one processor; and
[0040] A memory communicatively connected to the at least one processor; wherein,
[0041] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the word vector matrix enhancement method according to any embodiment of the present invention.
[0042] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the word vector matrix enhancement method according to any embodiment of the present invention when executed.
[0043] The technical solution of the embodiment of the present invention includes: obtaining a text set; determining an initial word vector matrix corresponding to the text set, where the word vector matrix is an M*N matrix, M is the number of word vectors, N is the dimension of the word vectors, and the dimension of the word vectors is equal to the number of texts; and performing enhancement processing on matrix elements with a quantization value of a first set value based on the initial word vector matrix to obtain a target word vector matrix. By means of relationship transmission, this technical solution performs data enhancement on the word vector matrix, so that word vectors composed of words with a relatively low frequency of occurrence can be better clustered.
[0044] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0046] Figure 1 is a flowchart of a word vector matrix enhancement method according to Embodiment 1 of the present invention;
[0047] Figure 2 is a schematic structural diagram of a word vector matrix enhancement device according to Embodiment 2 of the present invention;
[0048] Figure 3 is a schematic structural diagram of an electronic device according to Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0051] Embodiment 1
[0052] Figure 1 is a flowchart of a method for enhancing a word vector matrix according to Embodiment 1 of the present invention. This embodiment is applicable to the situation of enhancing a word vector matrix. This method can be executed by a word vector matrix enhancement device, which can be implemented in the form of hardware and / or software, and the word vector matrix enhancement device can be configured in an electronic device with data processing capabilities. As Figure 1 shown, the method includes:
[0053] S110. Obtain a text set.
[0054] Among them, the text set can be understood as a set composed of many texts. The text can contain various vocabulary. In this embodiment, various text sets can be obtained.
[0055] S120. Determine the initial word vector matrix corresponding to the text set.
[0056] Among them, the word vector matrix is a matrix of M*N. M can be the number of word vectors, N can be the dimension of the word vectors, and the dimension of the word vectors is equal to the number of texts. It can be understood that the dimension of the word vectors in this embodiment is equal to the number of texts in the text set.
[0057] In this embodiment, the semantic information of a word can be a set composed of feature information for describing the word. Among them, the feature information of a word can include but is not limited to at least one of the following: the meaning of the word, the part of speech (such as noun, adjective, etc.), synonyms, antonyms, etc. For example, the semantic information of "pleasant to the ear" can include: the meaning is "nice, a sound that can make people happy"; the synonym is "sweet-sounding"; the antonym is "harsh", etc. The semantic information of a vocabulary can include the feature information of each word included in the vocabulary. A word vector can be understood as a vector composed of numbers mapped by the feature information of a word. Each word corresponds to a unique word vector. A word vector matrix is a matrix composed of word vectors corresponding to each word included in the vocabulary. Usually, a row or a column of elements in the word vector matrix represents a word vector. In the embodiments of the present invention, if not otherwise specified, it is generally described by taking a row in the word vector matrix representing a word vector as an example. In this embodiment, an initial word vector matrix corresponding to the text set can be determined.
[0058] In this embodiment, optionally, determining the initial word vector matrix corresponding to the text set includes: obtaining a word to be processed; for the i-th word, if the i-th word does not appear in the j-th text, then determining the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as a first set value; if the i-th word appears in the j-th text, then determining the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as a second set value.
[0059] Among them, the word to be processed can be understood as a vocabulary that needs to be represented by a word vector. Word segmentation can be a process of recombining a continuous character sequence into a word sequence according to certain specifications. In this embodiment, the words to be processed can be obtained by dividing a document based on a word segmentation method, can be custom-defined words, can also be words obtained by dividing a document by setting delimiters, or can be words obtained by any hybrid method of the above three methods. In this embodiment, different quantization values can be set for the elements of the initialized matrix according to whether the word to be processed in the text appears in the text. The first set value can be understood as a pre-set value. In this embodiment, the first set value can be set to 0, or other representations of 0. The second set value can be understood as a pre-set value. The second set value is not equal to the first set value. In this embodiment, the second set value can be 1, or other values, which can be set according to actual needs.
[0060] In this embodiment, when obtaining the word to be processed, for the i-th word, if the i-th word has not been processed in the j-th text, then determining the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as the first set value 0; if the i-th word appears in the j-th text, then determining the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as the second set value 1.
[0061] In this embodiment, the word to be processed can be represented based on the vector representation method of the text. Exemplarily, assume that there are n texts in the text set. If certain words often appear in pairs in multiple identical texts, we consider that these two words are very closely related. For the text set, the texts can be numbered in sequence (i = 0...n - 1), and the text numbers are used as vector indices, so there is an n-dimensional vector. When a word appears in a certain text i, the value at vector position i can be 1, and in this way, a word can be represented by a vector in a form similar to [0, 1, 0,..., 1, 0].
[0062] Through such a setting in this embodiment, the word to be processed can be represented by converting it into a vector form. By setting different quantization values according to the characteristics of words co-occurring in the same document in the text, it is convenient to perform data enhancement on the word vector matrix through the way of relationship transmission.
[0063] S130. Perform enhancement processing on the matrix elements with the quantization value being the first set value based on the initial word vector matrix to obtain the target word vector matrix.
[0064] Among them, the target word vector matrix can be understood as being obtained by performing enhancement processing on the matrix elements with the quantization value being the first set value based on the initial word vector matrix. In this embodiment, enhancement processing can be performed on the matrix elements with the first set value being 0.
[0065] Exemplarily, in a text set, word A and word B co-occur in text D1, word B and word C co-occur in text D2, and word A and word C do not co-occur in any text. Then, data enhancement is performed on the word vectors of word A and word C, so as to increase the characteristics of word A and word C co-occurring in text D1 and text D2.
[0066] In this embodiment, optionally, performing enhancement processing on the matrix elements with the quantization value being the first set value based on the initial word vector matrix includes: if the quantization value of the element in the a-th row and b-th column of the initial word vector matrix is the first set value, then obtain N quantization values in the a-th word vector as the first quantization value; obtain the quantization values of M words corresponding to the b-th text as the second quantization value; and perform enhancement processing on the element in the a-th row and b-th column based on the first quantization value, the second quantization value, and the initial word vector matrix.
[0067] Among them, the first quantization value can be understood as N quantization values in the a-th word vector. N can be understood as the number of texts in the text set. The second quantization value can be understood as the quantization values of M words corresponding to the b-th text. M can be understood as the number of word vectors of the word to be processed. In this embodiment, enhancement processing can be performed on the element in the a-th row and b-th column based on the first quantization value, the second quantization value, and the initial word vector matrix.
[0068] In this embodiment, if the quantization value of the element in the a-th row and b-th column of the initial word vector matrix is the first set value, then obtain N quantization values in the a-th word vector as the first quantization value; the quantization values of M words corresponding to the b-th text can be obtained as the second quantization value, and then the element in the a-th row and b-th column is enhanced based on the first quantization value, the second quantization value, and the initial word vector matrix.
[0069] In this embodiment, through such a setting, the matrix elements with the first set value can be enhanced according to the quantization value, so that the word vectors composed of words with lower frequencies can be better clustered.
[0070] In this embodiment, optionally, enhancing the element in the a-th row and b-th column based on the first quantization value, the second quantization value, and the initial word vector matrix includes: multiplying the N first quantization values by the second set value and then summing to obtain a sum value; performing dot product of the a-th word vector with each word vector in the initial word vector matrix to obtain M dot product results; performing weighted summation on the M dot product results based on the M second quantization values; dividing the weighted summation result by the sum value to obtain the enhanced quantization value of the element in the a-th row and b-th column.
[0071] Among them, the sum value can be understood as the result of multiplying the N first quantization values by the second set value and then summing. The M dot product results can be understood as the results of performing dot product of the a-th word vector with each word vector in the initial word vector matrix.
[0072] In this embodiment, the N first quantization values can be summed to obtain a sum value; perform dot product of the a-th word vector with each word vector in the initial word vector matrix to obtain M dot product results; perform weighted summation on the M dot product results based on the M second quantization values, and then divide the weighted summation result by the sum value to obtain the enhanced quantization value of the element in the a-th row and b-th column.
[0073] In this embodiment, through such a setting, the quotient of the weighted summation result of the dot product results according to the second quantization value and the sum value obtained by summing the first quantization value can be obtained as the enhanced quantization value of the matrix element, which can improve the clustering effect of the word vector.
[0074] In this embodiment, optionally, the enhancement processing of the matrix element with the quantization value of the first set value based on the initial word vector matrix is calculated according to the following formula:
[0075]
[0076] Among them, M represents the number of word vector data to be processed, N represents the number of texts in the text set, Ha The word vector representing the a-th word, x ab represents the value at the b-th position in the word vector data of the a-th word; A represents a second set value. Among them, the value range of a can be 0 to M - 1; the value range of b can be 0 to N - 1.
[0077] In this embodiment, x ab represents the value at the b-th position in the word vector of the a-th word, x ab The value range of can be [0, 1]. If the word a exists in the b-th document, then x ab is initialized to 1, otherwise x ab is initialized to 0.
[0078] Exemplarily, Table 1 can be the initialized word vector matrix, and Table 2 can be the word vector matrix after data augmentation.
[0079] y1 y2 y3 y4 x1 1 1 1 0 x2 1 1 0 0 x3 0 0 1 1 x4 0 0 0 1
[0080] Table 1 Initialized word vector matrix
[0081] y1 y2 y3 y4 x1 1 1 1 1 / 3 x2 1 1 1 0 x3 1 / 2 1 / 2 1 1 x4 0 0 1 1
[0082] Table 2 Word vector matrix after data augmentation
[0083] Among them, as shown in Table 1, it is the word vector matrix M1 initialized and constructed for words x1, x2, x3, x4 in documents y1, y2, y3, y4. As shown in Table 2, it is the word vector matrix M2 obtained by performing data augmentation on the word vector matrix M1. Table 2 can be obtained by performing numerical calculations on the initial values according to the above data augmentation formula; among them, the initial values of the bolded values are 0, and the values of the non-bolded fonts are the same as the initial values.
[0084] For example: The word x1 appears in documents y1, y2, and y3, and the word x2 appears in documents y1 and y2. Since the words x1 and x2 co-occur in documents y1 and y2, and the word x2 only co-occurs in documents y1 and y2, and the second set value is 1, then the value x 23 in the word vector X2 of the word x2 is calculated through the formula, and the calculation process is as shown in Formula 2, and the obtained x 23 has a value of 1.
[0085]
[0086] The technical solution of the embodiment of the present invention includes: obtaining a text set; determining an initial word vector matrix corresponding to the text set, where the word vector matrix is an M*N matrix, M is the number of word vectors, N is the dimension of the word vectors, and the dimension of the word vectors is equal to the number of texts; and performing enhancement processing on matrix elements with a quantization value of a first set value based on the initial word vector matrix to obtain a target word vector matrix. Through the method of relationship transmission, this technical solution performs data enhancement on the word vector matrix, enabling better clustering analysis of word vectors composed of words with lower frequencies of occurrence.
[0087] Embodiment 2
[0088] Figure 2 is a schematic structural diagram of a word vector matrix enhancement device provided according to Embodiment 2 of the present invention. As Figure 2 shown, the device includes:
[0089] A text set acquisition module 210, configured to acquire a text set;
[0090] An initial word vector matrix determination module 220, configured to determine an initial word vector matrix corresponding to the text set, where the word vector matrix is an M*N matrix, M is the number of word vectors, N is the dimension of the word vectors, and the dimension of the word vectors is equal to the number of texts;
[0091] A target word vector matrix acquisition module 230, configured to perform enhancement processing on matrix elements with a quantization value of a first set value based on the initial word vector matrix to obtain a target word vector matrix.
[0092] Optionally, the initial word vector matrix determination module 220 includes:
[0093] A to-be-processed word acquisition unit, configured to acquire a to-be-processed word;
[0094] A first set value determination unit, configured to, for the i-th word, if the i-th word does not appear in the j-th text, determine the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as the first set value;
[0095] A second set value determination unit, configured to, if the i-th word appears in the j-th text, determine the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as the second set value.
[0096] Optionally, the target word vector matrix acquisition module includes:
[0097] A first quantization value acquisition unit, configured to, if the quantization value of the element in the a-th row and b-th column of the initial word vector matrix is the first set value, acquire N quantization values in the a-th word vector as the first quantization value;
[0098] A second quantization value acquisition unit, configured to acquire quantization values of M words corresponding to the b-th text as second quantization values;
[0099] An element enhancement processing unit, configured to perform enhancement processing on the element in the a-th row and b-th column based on the first quantization value, the second quantization value, and the initial word vector matrix.
[0100] Optionally, the element enhancement processing unit is configured to:
[0101] Multiply N first quantization values by a second set value and sum them to obtain a sum value;
[0102] Perform dot product of the a-th word vector with each word vector in the initial word vector matrix to obtain M dot product results;
[0103] Perform weighted summation on the M dot product results based on M second quantization values;
[0104] Divide the weighted summation result by the sum value to obtain the enhanced quantization value of the element in the a-th row and b-th column.
[0105] Optionally, the enhancement processing on the matrix element with a quantization value of a first set value based on the initial word vector matrix is calculated according to the following formula:
[0106]
[0107] where M represents the number of word vector data to be processed, N represents the number of texts in the text set, H a represents the word vector of the a-th word, and x ab represents the value at the b-th position in the word vector data of the a-th word; A represents the second set value.
[0108] A word vector matrix enhancement device provided by an embodiment of the present invention can execute a word vector matrix enhancement method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0109] Embodiment III
[0110] Figure 3It is a schematic structural diagram of an electronic device provided according to Embodiment 3 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0111] As Figure 3 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0112] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0113] The processor 11 can be various general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the word vector matrix enhancement method.
[0114] In some embodiments, the word vector matrix enhancement method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the word vector matrix enhancement method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the word vector matrix enhancement method by any other suitable means (e.g., by means of firmware).
[0115] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0116] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0117] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0118] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0119] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0120] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0121] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0122] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for enhancing a word vector matrix, characterized in that, Including: Obtain a text set; Determine an initial word vector matrix corresponding to the text set; wherein, the word vector matrix is a matrix of M*N, M is the number of word vectors, N is the dimension of the word vectors, and the dimension of the word vectors is equal to the number of texts; Perform enhancement processing on matrix elements with a quantization value of a first set value based on the initial word vector matrix to obtain a target word vector matrix; Determine an initial word vector matrix corresponding to the text set, including: Obtain a word to be processed; For the i-th word, if the i-th word does not appear in the j-th text, then determine the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as the first set value; If the i-th word appears in the j-th text, then determine the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as the second set value; Perform enhancement processing on matrix elements with a quantization value of a first set value based on the initial word vector matrix, including: If the quantization value of the element in the a-th row and b-th column of the initial word vector matrix is the first set value, then obtain N quantization values in the a-th word vector as the first quantization value; Obtain the quantization values of M words corresponding to the b-th text as the second quantization value; Perform enhancement processing on the element in the a-th row and b-th column based on the first quantization value, the second quantization value, and the initial word vector matrix; Perform enhancement processing on the element in the a-th row and b-th column based on the first quantization value, the second quantization value, and the initial word vector matrix, including: Multiply N first quantization values by the second set value and sum to obtain a sum value; Perform dot product of the a-th word vector with each word vector in the initial word vector matrix to obtain M dot product results; Perform weighted summation on the M dot product results based on M second quantization values; Divide the weighted summation result by the sum value to obtain the enhanced quantization value of the element in the a-th row and b-th column; The enhancement processing of matrix elements with a quantization value of a first set value based on the initial word vector matrix is calculated according to the following formula: Among them, M represents the number of word vector data to be processed, N represents the number of texts in the text set, and H a represents the word vector of the a-th word, and x ab represents the value at the b-th position in the word vector data of the a-th word; A represents the second set value.
2. A word vector matrix enhancement device, characterized in that Including: A text set acquisition module for obtaining a text set; An initial word vector matrix determination module for determining an initial word vector matrix corresponding to the text set; wherein, the word vector matrix is a matrix of M*N, M is the number of word vectors, N is the dimension of the word vectors, and the dimension of the word vectors is equal to the number of texts; A target word vector matrix acquisition module for performing enhancement processing on matrix elements with a quantization value of a first set value based on the initial word vector matrix to obtain a target word vector matrix; The initial word vector matrix determination module includes: A word to be processed acquisition unit for obtaining a word to be processed; A first set value determination unit for, for the i-th word, if the i-th word does not appear in the j-th text, then determining the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as the first set value; A second set value determination unit for, if the i-th word appears in the j-th text, then determining the quantization value of the element in the i-th row and j-th column of the initial word vector matrix as the second set value; The target word vector matrix acquisition module includes: The first quantization value acquisition unit is configured to, if the quantization value of the element in the \(a\)-th row and \(b\)-th column of the initial word vector matrix is the first set value, acquire \(N\) quantization values in the \(a\)-th word vector as the first quantization values; The second quantization value acquisition unit is configured to acquire the quantization values of \(M\) words corresponding to the \(b\)-th text as the second quantization values; The element enhancement processing unit is configured to perform enhancement processing on the element in the \(a\)-th row and \(b\)-th column based on the first quantization values, the second quantization values, and the initial word vector matrix; The element enhancement processing unit is configured to: Multiply the \(N\) first quantization values by the second set value and sum them to obtain a sum value; Perform dot multiplication of the \(a\)-th word vector with each word vector in the initial word vector matrix to obtain \(M\) dot multiplication results; Perform weighted summation on the \(M\) dot multiplication results based on the \(M\) second quantization values; Divide the weighted summation result by the sum value to obtain the quantization value of the enhanced element in the \(a\)-th row and \(b\)-th column; The enhancement processing of the matrix element with the quantization value of the first set value based on the initial word vector matrix is calculated according to the following formula: Among them, M represents the number of word vector data to be processed, N represents the number of texts in the text set, and H a represents the word vector of the a-th word, and x ab represents the value at the b-th position in the word vector data of the a-th word; A represents the second set value.
3. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the word vector matrix enhancement method according to any one of claims 1.
4. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the word vector matrix enhancement method according to any one of claims 1 when executed.
Citation Information
Patent Citations
Semantic-enhanced relation extraction method and device, computer equipment and storage medium
CN113626608A
Text data enhancement method and device, electronic equipment and storage medium
CN114298024A