Word Vector Generation Method, Device, Equipment and Medium Incorporating Word Granularity Information
By integrating word and character-level information through linear transformation and an ESIMCSE model, the method enhances sentence vector generation accuracy and efficiency, addressing the limitations of existing methods in complex language environments.
Patent Information
- Application Number
- CN202211097605.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-09-08
AI Technical Summary
When processing Chinese sentences, the existing sentence vector generation method ignores the word order information between words and words, resulting in the generated sentence vector data that cannot accurately represent the semantics of the sentence, especially in complex language environments. The method based on word granularity has incomplete training and Oov problems, while the method based on word granularity is not very differentiated.
By obtaining training statement data, word vector data and word vector data are obtained, and linear layer mapping is performed, so that the word vector data dimensions are the same as the word vector data dimensions, the associated word vector data and word vector data are added, and the unrelated word vector data is recorded as separate data, and the sentence vector data is input to the unsupervised training model to generate sentence vector data.
The generated sentence vector data is efficient and convenient while maintaining the same meaning, improving the accuracy and distinction of sentence vector data, and is suitable for complex locale environments.
Smart Images

Figure CN116187300B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of natural language processing and intelligent question answering, and specifically relates to a method, device, equipment and medium for generating sentence vectors by integrating word and character granularity information. Background Art
[0002] The generation of sentence vector data is a basic task in the field of natural language processing and plays an important role in downstream tasks such as text classification, clustering, and similarity calculation. Therefore, the method for generating sentence vector data has always been a hot topic of research. Among the existing methods for generating sentence vector data, some are based on the bag-of-words method, which simply tokenizes sentences or divides them into individual characters, and then takes the average of all characters or words to obtain sentence vector data. Such a generation method ignores the word order information between words and loses the semantic information of sentences, resulting in the generated sentence vector data being unable to accurately represent sentences. In addition, there are also ways to generate sentence vector data through models, which obtain sentence vector data by dynamically encoding based on word vector data and can better represent the semantic information of sentences. However, for some complex language environments, the existing model-based methods for generating sentence vector data still cannot achieve good results.
[0003] Currently, some Chinese sentence vector data generation methods are based on word granularity, and some are based on character granularity. If based on word granularity, the number of Chinese words is extremely large, on the order of tens of millions or even hundreds of millions, and there are a large number of long-tail words. Since these words appear infrequently, it is easy to have the problem of incomplete training. If the size of the word list is reduced, it is very easy to have the situation of oov (out of vocabulary, words outside the word list), which will affect the sentence vector data, and the effect of generating sentence vector data is extremely susceptible to the effect of word segmentation. If based on character granularity, the advantage is that the character list is not as large as the word list, and it is easy to find sufficient corpus for training. However, for sentence pairs that differ by one or two characters but have completely different meanings, the discrimination of the existing character granularity-based sentence vector generation method is not high. When calculating the cosine similarity using the generated sentence vectors, the similarity will be very high. For example: Add WeChat; Add WeChat group. There are still areas for improvement in the current methods for generating sentence vector data. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, the present invention provides a method, device, equipment and medium for generating sentence vectors by integrating word and character granularity information to solve the technical problems existing in the above-mentioned sentence vector generation method, and the following technical solutions can be proposed.
[0005] The present invention proposes a method for generating sentence vectors by integrating word and character granularity information, and the method includes:
[0006] Obtain training statement data, and perform word segmentation on the training statement data to obtain word data;
[0007] Obtain text input data, and parse the text input data to obtain word vector data, character vector data, position vector data corresponding to the word vector data and the character vector data, and character vector data corresponding to the word vector data and the character vector data;
[0008] Perform linear layer mapping processing on the word vector data to change the vector data dimension of the word vector data so that the vector data dimension of the word vector data is the same as the vector data dimension of the character vector data;
[0009] According to the word data, add the associated word vector data and character vector data, record the word vector data and the word data as extended data, and record the character vector data as the first separate data; according to the word data, record the unassociated character vector data as the second separate data;
[0010] Input the character vector data, the word vector data, the position vector data, the character vector data, the extended data, the first separate data, and the second separate data into an unsupervised training model to generate sentence vector data of the text input data.
[0011] In an embodiment of the present invention, after the step of inputting the character vector data, the word vector data, the position vector data, the character vector data, the extended data, the first separate data, and the second separate data into an unsupervised training model to generate sentence vector data of the text input data, it includes:
[0012] Set the character identifier of the extended data as the first character identifier, and set the character identifiers of the first separate data and the second separate data as the second character identifier;
[0013] Detect the accuracy of the sentence vector data of the text input data according to the first character identifier and the second character identifier.
[0014] In an embodiment of the present invention, the unsupervised training model is an esimcse unsupervised training model, which is trained with the character vector data, the word vector data, the position vector data, the character vector data, the extended data, the first separate data, and the second separate data in all text input data.
[0015] In an embodiment of the present invention, the loss function of the esimcse unsupervised training model is as follows:
[0016]
[0017] Where li denotes the loss of the i-th sample in a batch process, N denotes the size of the batch process, sim() represents the calculation using cosine similarity, and h i denotes the sentence vector data of the i-th sample in a batch process, h i + denotes the sentence vector data of the positive example of the i-th sample in a batch process, h j + denotes the sentence vector data of the positive example of the j-th sample in this batch process, h m + denotes the sentence vector data of the m-th sample in the momentum queue, and τ represents the temperature hyperparameter.
[0018] In an embodiment of the present invention, the step of obtaining training sentence data and performing word segmentation on the training sentence data to obtain word data includes:
[0019] Obtain training sentence data, perform word segmentation on the training sentence data to obtain word data, and count the word frequency of the word data that appears within a unit time;
[0020] Preset a word frequency threshold, obtain the word data that exceeds the word frequency threshold, and construct a word library.
[0021] In an embodiment of the present invention, the step of presetting a word frequency threshold, obtaining the word data that exceeds the word frequency threshold, and constructing a word library includes:
[0022] Preset a word frequency threshold, obtain the word data that exceeds the word frequency threshold, and assign weights to the word data according to the word frequency of the word data;
[0023] Construct a word library according to the weights of the word data.
[0024] In an embodiment of the present invention, after the step of presetting a word frequency threshold, obtaining the word data that exceeds the word frequency threshold, and assigning weights to the word data according to the word frequency of the word data, it includes:
[0025] Judge whether a word data participates in the composition of other word data. When a word data participates in the composition of other word data, assign the same weight to the word data and other said word data;
[0026] When a word data does not participate in the composition of other word data, assign a weight to the word data;
[0027] Construct a word library according to the weights of the word data.
[0028] The present invention proposes a sentence vector generation device incorporating word and character granularity information, and the device includes:
[0029] An acquisition unit, configured to acquire training statement data, and perform word segmentation processing on the training statement data to obtain word data;
[0030] An analysis unit, configured to acquire text input data, and analyze the text input data to obtain word vector data, character vector data, position vector data corresponding to the word vector data and the character vector data, and character vector data corresponding to the word vector data and the character vector data;
[0031] A transformation unit, configured to perform linear layer mapping processing on the word vector data, and transform the vector data dimension of the word vector data, so that the vector data dimension of the word vector data is the same as the vector data dimension of the character vector data;
[0032] A recording unit, configured to add the associated word vector data and character vector data according to the word data, record the word vector data and the word data as extended data, record the character vector data as first separate data; according to the word data, record the unassociated character vector data as second separate data;
[0033] A generation unit, configured to input the character vector data, the word vector data, the position vector data, the character vector data, the extended data, the first separate data, and the second separate data into an unsupervised training model to generate sentence vector data of the text input data.
[0034] The present invention provides an electronic device, which includes:
[0035] One or more processors;
[0036] A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the electronic device to implement the sentence vector generation method incorporating word and character granularity information as described in any one of the above.
[0037] The present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor of a computer, enable the computer to execute the sentence vector generation method incorporating word and character granularity information as described in any one of the above.
[0038] Advantageous effects of the present invention: The present invention provides a sentence vector generation method, device, equipment and medium incorporating word and character granularity information. After adding the word vector data and the character vector data, while ensuring that the generated sentence vector data has the same meaning as the text input data, it also has high efficiency and convenience.
[0039] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit this application. Description of the Drawings
[0040] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other accompanying drawings based on these drawings without creative efforts. In the accompanying drawings:
[0041] Figure 1 It is a schematic diagram of the implementation environment of the sentence vector generation method incorporating word and character granularity information shown in an exemplary embodiment of this application.
[0042] Figure 2 It is a flowchart of the sentence vector generation method incorporating word and character granularity information shown in an exemplary embodiment of this application.
[0043] Figure 3 is Figure 2 It is a flowchart of step S290 in the shown embodiment in an exemplary embodiment.
[0044] Figure 4 is Figure 2 It is a flowchart of step S210 in the shown embodiment in an exemplary embodiment.
[0045] Figure 5 It is a schematic framework diagram of the sentence vector generation method incorporating word and character granularity information shown in an exemplary embodiment of this application.
[0046] Figure 6 It is a schematic model diagram of the sentence vector generation method incorporating word and character granularity information shown in an exemplary embodiment of this application.
[0047] Figure 7 is Figure 6 It is a schematic model diagram of module 603 in the shown embodiment in an exemplary embodiment.
[0048] Figure 8 is Figure 6 It is a schematic structural diagram of module 603 in the shown embodiment in an exemplary embodiment.
[0049] Figure 9 It is a schematic structural diagram of the sentence vector generation device incorporating word and character granularity information shown in an exemplary embodiment of this application.
[0050] Figure 10 It is a schematic structural diagram of a computer device shown in an exemplary embodiment of this application. Detailed implementation manners
[0051] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for explaining the present invention, rather than limiting the protection scope of the present invention.
[0052] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0053] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0054] First of all, it should be noted that the concept of sentence vector data is similar to that of word vector data, which is to project the sentence semantics onto an n-dimensional vector data space. The application scenarios of sentence vector data generally include semantic retrieval, text clustering, and text classification. In addition to these direct application scenarios, in other NLP (Natural Language Processing) tasks, the quality of the intermediate product sentence vector data will largely affect the quality of the task results. For example, in the seq2seq task, the intermediate semantic vector data c, and in the long text NLP task, attention is performed on multiple sentence vector data to select important sentences, which is also very important for grasping the overall long text NLU (Natural Language Understanding) and NLG (Natural Language Generation). Among them, seq2seq is sequence-to-sequence, which generates one sequence from another sequence. It involves two processes: one is to understand the previous sequence, and the other is to generate a new sequence with the understood content. The attention mechanism refers to the attention mechanism in the field of deep learning.
[0055] In addition, granularity measures the amount of information contained in a text. If a text contains a large amount of information, its granularity is large; conversely, if it contains little information, its granularity is small. With this principle, it is easy for us to judge the granularity of a text. Words like "lingering", "rugged", and "grape", although composed of two characters, express only one meaning, so the granularity of these words is small. Words like "basketball" and "mouse pad" are composed of simple words. Although they also express only one meaning, they can be split, such as "basket" and "ball", "mouse" and "pad". The granularity of such words is slightly larger. Words like "laptop" and "HD set-top box" have an even larger granularity. Proper names are a special type of words. Although they contain many characters, they actually express only one meaning. For example, the names of movies and TV shows like "Scarlet Heart: Return to Zero" and "Home with a Fang" have a very small granularity. Organization names, personal names, etc. are proper names with internal structures and have a slightly larger granularity than movie names. Obviously, when discussing text granularity, the ideal way is to start from the semantic perspective and conduct reasonable analysis and judgment.
[0056] Figure 1 It is a schematic diagram of the implementation environment of the sentence vector generation method incorporating word granularity information shown in an exemplary embodiment of the present application. As Figure 1 shown, in some embodiments, the current user of the client 110 can send an input instruction to the server 130 through a communication network. After receiving the input instruction from the client 110, the server 130 can process the sentence vector data based on word granularity. Among them, Figure 1 the server 110 shown can be any terminal device that supports the installation of navigation map software, such as a smart phone, in-vehicle computer, tablet computer, laptop computer, or wearable device, but is not limited thereto. Figure 1 The server 220 shown is a server. For example, it can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. There is no limitation here either. The client 110 can communicate with the server 130 through wireless networks such as 3G (the third-generation mobile information technology), 4G (the fourth-generation mobile information technology), and 5G (the fifth-generation mobile information technology). There is no limitation here either.
[0057] Figure 2 It is a schematic diagram of the implementation environment for refreshing road conditions during navigation shown in an exemplary embodiment of the present application. In some embodiments, this method can be applied to Figure 1The illustrated implementation environment is specifically executed by the client 110 in this implementation environment. It should be understood that this method can also be applied to other exemplary implementation environments and be specifically executed by devices in other implementation environments. This embodiment does not limit the implementation environment applicable to this method.
[0058] In some embodiments, by way of example, in the client 110 to which the method for generating sentence vectors incorporating word-level granularity information disclosed in this embodiment is applicable, an SDK (Software Development Kit, a collection of development tools for building application software for specific software packages, software frameworks, operating systems, etc.) can be installed, and the method disclosed in this embodiment is specifically implemented as one or more functions provided by the SDK.
[0059] Such as Figure 2 As shown, in an exemplary embodiment, a method for generating sentence vectors incorporating word-level granularity information includes at least steps S210 to S290, which are introduced in detail as follows:
[0060] Step S210, obtain training statement data, and perform word segmentation on the training statement data to obtain word data.
[0061] In some embodiments, first, training corpus can be obtained. The training corpus can be specifically set in combination with sentence vector data. For example, in the semiconductor field, semiconductor-related statements can be obtained for arranging the statement environment. The training corpus can be segmented using a word segmentation tool to obtain word data. For example, jieba or fastNLP can be used to segment the corpus. Word segmentation refers to the process of splitting a sequence of Chinese characters into individual words, that is, the process of recombining a continuous sequence of characters into a sequence of words according to certain specifications. After performing word segmentation on the training corpus, the word data can be statistically analyzed, and the importance degree of the word data can be reflected by the quantity of the word data. The word data can be statistically analyzed and a word library can be constructed.
[0062] Step S230, obtain text input data, and parse the text input data to obtain word vector data, word vector data, position vector data corresponding to the word vector data and the word vector data, and character vector data corresponding to the word vector data and the word vector data.
[0063] In some embodiments, text input data input by the current user can be obtained, the current text input data can be parsed, and word vector data, word vector data, position vector data corresponding to the word vector data and character vector data corresponding to the word vector data can be obtained. The Chinese natural language processing vector data collection includes word vector data, pinyin vector data, word vector data, part-of-speech vector data, and dependency relationship vector data. There are a total of five types of vector data. Among them, word vector data (Word embedding) is a general term for a set of language modeling and feature learning techniques in Word-embedded natural language processing (NLP), where words or phrases from a vocabulary are mapped to vector data of real numbers. Conceptually, it involves a mathematical embedding from the one-dimensional space of each word to a continuous vector data space with a lower dimension. The concept of word vector data is similar to that of word vector data, that is, the word semantics is projected onto an n-dimensional vector data space, and corresponding word vector data can be formed between specific word vector data. A character vector data is a vector data composed of strings. A character here does not refer to a single letter or symbol in the literary sense, but a string like this is astring. Both double quotes and single quotes can be used to generate character vector data. The position vector data reflects the position information of the word vector data and word vector data in the text input data, and the position information in the sentence vector data can be determined through the position vector data.
[0064] Step S250, perform a linear layer mapping process on the word vector data to change the vector data dimension of the word vector data so that the vector data dimension of the word vector data is the same as the vector data dimension of the word vector data.
[0065] In some embodiments, the word vector data can be subjected to a linear layer mapping process. In the linear layer mapping process, combining additivity and proportionality together is the core meaning of "linear". For example, write the values that the independent variable and dependent variable can take on a one-dimensional coordinate axis, and understand each value on the coordinate axis as a vector data from the origin to the corresponding point of the value, which is mapped as a mapping from vector data to vector data. After the word vector data is subjected to the linear layer mapping process, the vector data dimension of the word vector data can be changed. The vector data dimension refers to the number of samples or the number of features. Generally, without special instructions, it refers to the number of features. For example, for an array, the vector data dimension is the result returned by the function shape. The number of numbers returned in shape is the number of dimensions. For an image, the vector data dimension is the number of feature vector data in the image. Make the vector data dimension of the word vector data the same as the vector data dimension of the word vector data, so that the word vector data and word vector data can be added.
[0066] Step S270, according to the word data, add the associated word vector data and character vector data, record the word vector data and the word data as extended data, record the character vector data as first separate data; according to the word data, record the unassociated character vector data as second separate data.
[0067] In some embodiments, according to the word data in the vocabulary, the associated word vector data and the word vector data are added. For example, the word vector data and the word vector data can be added so that the word vector data and the word vector data are the same as the word data after addition. After the word vector data and the word vector data are added, the word vector data and the word data can be recorded as extended data, which can indicate that the word vector data is contained in the word data. For example, in the sentence "Coach, I want to play rugby", the word data can be "rugby" and the word vector data can be "olive", "coach". The word vector data is recorded as the first individual data. For example, the first individual data can be "olive", "rugby", "ball". In addition, the unassociated word vector data is recorded as the second individual data. For example, in the sentence "Coach, I want to play rugby", the second individual data can be "I", "want", "play".
[0068] Step S290, input the word vector data, the word vector data, the position vector data, the character vector data, the extended data, the first individual data, and the second individual data into an unsupervised training model to generate sentence vector data for the text input data.
[0069] In some embodiments, the word vector data, the word vector data, the position vector data, the character vector data, the extended data, the first separate data, and the second separate data can be input into an unsupervised training model to generate sentence vector data for the text input data. Among them, the unsupervised training model can be an esimcse unsupervised training model. Esmicse can construct positive sample pairs by "word repetition" and introduce a momentum comparison mechanism to increase the number of negative samples during contrast learning without increasing batchSize (batch processing), so that the pairings can be more fully compared during model training. When constructing the model input, labeled data will still be used, but there is no need to perform the "word repetition" step, but directly construct positive sample pairs.
[0070] The unsupervised training model is the esimcse unsupervised training model, which is trained with the word vector data, word vector data, position vector data, character vector data, extended data, first individual data, and second individual data in all text input data. And the loss function of the esimcse unsupervised training model during training is as follows:
[0071]
[0072] where l i represents the loss of the i-th sample in a batch process, N represents the size of the batch process, sim() represents the use of cosine similarity calculation, and h i represents the sentence vector data of the i-th sample in a batch process, and h i + represents the sentence vector data of the positive example of the i-th sample in a batch process, and h j + represents the sentence vector data of the positive example of the j-th sample in this batch process, and h m + represents the sentence vector data of the m-th sample in the momentum queue, and τ represents the temperature hyperparameter.
[0073] Please refer to Figure 3 as shown Figure 3 is Figure 2 a flowchart of step S290 in the illustrated embodiment in an exemplary embodiment. The sentence vector generation method incorporating word and character granularity information may include steps S310 to S330, which are introduced in detail as follows.
[0074] In some embodiments, step S310 can be executed to set the character identifier of the extended data as the first character identifier, and set the character identifiers of the first individual data and the second individual data as the second character identifier. After step S310 is executed, step S330 can be executed. In step S330, according to the first character identifier and the second character identifier, the accuracy of the sentence vector data of the text input data is detected. The associated word vector data and word data can be queried through the first character identifier, and whether the sentence vector data contains the word vector data and the word data can be queried through the word vector data and the word data. When the sentence vector data contains the word vector data and the word data, it indicates that the sentence vector data generation process is accurate. When the sentence vector data does not contain the word vector data and the word data, it indicates that the sentence vector data generation process is inaccurate.
[0075] Please refer to Figure 4 as shown Figure 4 is Figure 2 a flowchart of step S210 in the illustrated embodiment in an exemplary embodiment. The sentence vector generation method incorporating word and character granularity information may include steps S401 to S406, which are introduced in detail as follows.
[0076] In some embodiments, step S401 may be executed first. In step S401, training statement data may be obtained first, the training statement data may be segmented to obtain word data, and the word frequency of the word data appearing within a unit time may be counted. For example, in the sentence "Coach, I want to play rugby", the word frequency of words such as "coach", "olive", and "rugby" may be counted. After step S401 is executed, step S402 may be executed. In step S402, a preset word frequency threshold is set, word data exceeding the word frequency threshold is obtained, and weights are assigned to the word data according to the word frequency of the word data. For example, in the sentence "Coach, I want to play rugby", the preset word frequency threshold is 50. When the occurrence word frequency of "coach" is 100, the occurrence word frequency of "olive" is 150, and the occurrence word frequency of "rugby" is 60, the weights of "olive", "coach", and "rugby" may be assigned different weights in descending order of word frequency. Additionally, since both "olive" and "rugby" contain "olive", the same weight may be assigned to "olive" and "rugby". In step S403, it is determined whether a word data participates in the composition of other word data. When a word data contains other word data, the same weight may be assigned to the word data and the other word data. For example, the word data of "rugby" contains the word data of "olive", and the same weight may be assigned to "rugby" and "olive". When a word data does not contain other word data, a weight is assigned to the word data. Finally, step S406 may be executed to construct a word library according to the weights of the word data.
[0077] Please refer to Figure 5 as shown Figure 5 is a schematic framework diagram of a sentence vector generation method incorporating word and character granularity information shown in an exemplary embodiment of the present application. In some embodiments, training samples 520 may be extracted from the dataset 510. In the training samples 520, the training statement data may be segmented to obtain word data. The dataset 510 also includes text input data. The text input data may be parsed to obtain data information such as word vector data, character vector data, position vector data of word vector data and character vector data, and character vector data of word vector data and character vector data. According to the data information of the word vector data, character vector data, position vector data, and character vector data, operations of word vector data pre-training 530 and sentence vector data samples 540 to be generated may be performed. The word vector data pre-training 530 may train the data structure, data information, vector data dimension, etc. of the word vector data so that the word vector data can be added to other character vector data. In the sentence vector data samples 540 to be generated, the sentence vector data may be pre-processed, and various possible generation schemes of the sentence vector data may be given in advance. After being input into an unsupervised training model, the sentence vector data of the text input data may be generated.
[0078] Please refer toFigure 6 , Figure 7 and Figure 8 as shown in Figure 6 is a schematic diagram of a model of a sentence vector generation method incorporating word granularity information shown in an exemplary embodiment of the present application. In some embodiments, a model derived from a Transformer encoder (the encoder of the feature extractor) can be used, which has powerful encoding capabilities. The present application introduces word granularity information on the basis of the character granularity based on the self-attention mechanism to assist the model in context encoding. First, the text input needs to be converted into a vector data input, and the specific conversion process is as Figure 7 shown. The position embedding (position vector data) and the segment embedding (character vector data) remain unchanged and are consistent with the input information of the text input data. The word embedding (word vector data) is obtained by adding two parts of vector data, namely the character vector data and the word vector data. Among them, the character vector data is obtained from the token input ids (character vector data address 701) vector data, and the word vector data is obtained from the word input ids (word vector data address 702) vector data. The word vector data will undergo a linear layer mapping to change the vector data dimension to make it consistent with the character vector data dimension, so that it can be added to the character vector data. In the first layer of the transformer encoder (the encoder of the feature extractor) of the model, the word attention mask is used to introduce word information, and the subsequent transformer encoders use the token attention mask (self-attention mechanism) for self-attention calculation. As Figure 8As shown in the figure, taking "Coach, I want to play rugby" as an example, we can see that the word vector data and word vector data are first concatenated to calculate the token attention mask and word attention mask. The length of the token attention mask and word attention mask is equal to the length of the word vector data and the word vector data after concatenation, both of which are 16 character identifiers, and the token attention mask and word attention mask of each word vector data are recorded as 1. If a word vector data does not participate in the composition of a word vector data, then in the detection result of the word vector data, the word attention mask of the word vector data is equal to the token attention mask and both are recorded as 0. If a word vector data participates in the composition of a word vector data, it is calculated with the word vector data, so the word attention mask of the word vector data is 1 and the token attention mask is 0. For example, the word vector data of "olivine" in the example participates in the composition of the two word vector data of "olive" and "rugby", then the word attention mask of the word vector data and the two word vector data are both 1.
[0079] Figure 9 is a block diagram of a sentence vector generation device incorporating word granularity information, shown in an exemplary embodiment of the present application. The device can be applied to Figure 1 The implementation environment shown is specifically configured in the intelligent terminal 110. The device may also be applicable to other exemplary implementation environments and specifically configured in other devices. This embodiment does not limit the implementation environment applicable to the device.
[0080] like Figure 9As shown in the figure, the exemplary sentence vector generation device incorporating word and character granularity information includes an acquisition unit 901, a parsing unit 902, a transformation unit 903, a recording unit 904, and a generation unit 905. Among them, the acquisition unit 901 is used to acquire training sentence data and perform word segmentation processing on the training sentence data to obtain word data. The parsing unit 902 is used to acquire text input data and parse the text input data to obtain word vector data, character vector data, position vector data corresponding to the word vector data and the character vector data, and character vector data corresponding to the word vector data and the character vector data. The transformation unit 903 is used to perform linear layer mapping processing on the word vector data, transform the vector data dimension of the word vector data, and make the vector data dimension of the word vector data the same as the vector data dimension of the character vector data. The recording unit 904 is used to add the associated word vector data and character vector data according to the word data, record the word vector data and the word data as extended data, record the character vector data as the first separate data, and record the unassociated character vector data as the second separate data according to the word data. The generation unit 905 is used to input the character vector data, the word vector data, the position vector data, the character vector data, the extended data, the first separate data, and the second separate data into an unsupervised training model to generate sentence vector data of the text input data.
[0081] It should be noted that the sentence vector generation device incorporating word and character granularity information provided in the above embodiment belongs to the same concept as the sentence vector generation method incorporating word and character granularity information provided in the above embodiment. The specific ways in which each module and unit perform operations have been described in detail in the method embodiment, and will not be elaborated here. In practical applications, the road condition refresh device provided in the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here either.
[0082] An embodiment of the present application also provides an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the sentence vector generation method incorporating word and character granularity information provided in each of the above embodiments.
[0083] Figure 10 The structure diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. It should be noted that Figure 10 The computer system 1000 of the electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present application.
[0084] AsFigure 10 As shown in Figure 10 , the computer system 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes according to a program stored in a Read-Only Memory (ROM) 1002 or a program loaded from a storage section 1008 into a Random Access Memory (RAM) 1003, such as executing the methods described in the above embodiments. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.
[0085] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage section 1008 as needed.
[0086] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by a Central Processing Unit (CPU) 1001, various functions defined in the system of the present application are executed.
[0087] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0089] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the units themselves.
[0090] Another aspect of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of the computer, the computer is caused to execute the sentence vector generation method incorporating word and character granularity information as described above. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist alone without being assembled into the electronic device.
[0091] Another aspect of this application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the sentence vector generation method incorporating word and character granularity information provided in the above various embodiments.
[0092] The above embodiments are only used to exemplarily illustrate the principles and effects of the present invention, rather than to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A sentence vector generation method incorporating word-level granularity information, characterized in that The method includes: Obtain training statement data, and perform word segmentation on the training statement data to obtain word data; Obtain text input data, and parse the text input data to obtain word vector data, character vector data, corresponding position vector data between the word vector data and the character vector data, and character vector data corresponding to the word vector data and the character vector data; Perform linear layer mapping processing on the word vector data to change the vector data dimension of the word vector data, so that the vector data dimension of the word vector data is the same as the vector data dimension of the character vector data; According to the word data, add the associated word vector data and character vector data, record the word vector data and the word data as extended data, and record the character vector data as the first separate data; according to the word data, record the unassociated character vector data as the second separate data; Input the character vector data, the word vector data, the position vector data, the character vector data, the extended data, the first separate data, and the second separate data into an unsupervised training model to generate sentence vector data of the text input data.
2. The sentence vector generation method incorporating word particle size information according to claim 1, wherein After the step of inputting the character vector data, the word vector data, the position vector data, the character vector data, the extended data, the first separate data, and the second separate data into an unsupervised training model to generate sentence vector data of the text input data, it includes: Set the character identifier of the extended data as the first character identifier, and set the character identifiers of the first separate data and the second separate data as the second character identifier; Detect the accuracy of the sentence vector data of the text input data according to the first character identifier and the second character identifier.
3. The sentence vector generation method incorporating word granularity information according to claim 1, characterized in that The unsupervised training model is an esimcse unsupervised training model, which is trained with the character vector data, the word vector data, the position vector data, the character vector data, the extended data, the first separate data, and the second separate data in all text input data.
4. The sentence vector generation method incorporating word granularity information according to claim 3, wherein The loss function of the esimcse unsupervised training model is as follows: where l i represents the loss of the i-th sample in a batch process, N represents the size of the batch process, sim() represents the use of cosine similarity calculation, and h i represents the sentence vector data of the i-th sample in a batch process, and h i + represents the sentence vector data of the positive example of the i-th sample in a batch process, and h j + represents the sentence vector data of the positive example of the j-th sample in this batch process, and h m + represents the sentence vector data of the m-th sample in the momentum queue, and τ represents the temperature hyperparameter.
5. The sentence vector generation method incorporating word granularity information according to claim 1, characterized in that, The step of obtaining training statement data and performing word segmentation on the training statement data to obtain word data includes: Obtain training statement data, perform word segmentation on the training statement data to obtain word data, and count the word frequency of the word data appearing within a unit time; Preset a word frequency threshold, obtain the word data exceeding the word frequency threshold, and construct a word library.
6. The sentence vector generation method incorporating word granularity information according to claim 5, characterized in that The step of presetting a word frequency threshold, obtaining the word data exceeding the word frequency threshold, and constructing a word library includes: Preset a word frequency threshold, obtain the word data exceeding the word frequency threshold, and assign weights to the word data according to the word frequency of the word data; Construct a word library according to the weights of the word data.
7. The sentence vector generation method incorporating word granularity information according to claim 6, characterized in that, After the step of presetting a word frequency threshold, obtaining the word data exceeding the word frequency threshold, and assigning weights to the word data according to the word frequency of the word data, it includes: Determine whether a word data participates in the composition of other word data. When a word data participates in the composition of other word data, assign the same weight to the word data and the other word data; When a word data does not participate in the composition of other word data, assign a weight to the word data; Construct a word library according to the weights of the word data.
8. A sentence vector generation device incorporating word - character - level information, characterized in that, The device includes: An acquisition unit, configured to acquire training statement data, and perform word segmentation processing on the training statement data to obtain word data; An analysis unit, configured to acquire text input data, and analyze the text input data to obtain word vector data, character vector data, position vector data corresponding to the word vector data and the character vector data, and character vector data corresponding to the word vector data and the character vector data; A transformation unit, configured to perform linear layer mapping processing on the word vector data, and transform the vector data dimension of the word vector data to make the vector data dimension of the word vector data the same as the vector data dimension of the character vector data; A recording unit, configured to add the associated word vector data and character vector data according to the word data, record the word vector data and the word data as extended data, and record the character vector data as first separate data; according to the word data, record the unassociated character vector data as second separate data; A generation unit, configured to input the character vector data, the word vector data, the position vector data, the character vector data, the extended data, the first separate data, and the second separate data into an unsupervised training model to generate sentence vector data of the text input data.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the electronic device to implement the sentence vector generation method incorporating word and character granularity information according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor of a computer, enable the computer to execute the sentence vector generation method incorporating word and character granularity information according to any one of claims 1 to 7.