Word Embedding Representation Learning Method and Apparatus, Text Recall Method and Apparatus
By constructing graph structure and random walk to obtain node sequences, the word embedding representation model is trained, which solves the recall errors or incompleteness caused by typos, and improves the accuracy and integrity of the recall.
Patent Information
- Application Number
- CN202010961808.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-09-14
AI Technical Summary
During the information recall process, the search string entered by the user may have typos, resulting in incorrect or incomplete recall results, which reduces the user experience.
By constructing a graph structure, using word segmentation and pronunciation information to construct node relationships, randomly access node sequences, and training the word embedding representation model based on these sequences to generate word embedding lookup tables, so that words with similar morphology have similar distances in vector space.
Improves the accuracy and integrity of the recall, and can recall information containing search strings and similar words to the lexical morphology, avoiding recall errors or missing caused by input errors.
Smart Images

Figure CN112100332B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of natural language processing. Specifically, it relates to a word embedding representation learning method, a word embedding representation learning device, a text recall method, a text recall device, a computer-readable storage medium, and an electronic device. Background Art
[0002] Word embedding, also known as word vector, word representation, text representation, etc., is a general term for language model and representation learning techniques in natural language processing (NLP). It refers to embedding a high-dimensional space with a dimension equal to the number of all words into a much lower-dimensional continuous vector space, and each word or phrase is mapped to a vector in the real number domain.
[0003] When recalling information according to a search string, the user may inadvertently have misspelled characters in the search string. For example, the search string the user wants to input is "COVID-19", but the actual input search string is "new official pneumonia". If the recall is strictly performed according to the search string containing misspelled characters, there will be cases where the recall result is incorrect or incomplete, lacking the recall result corresponding to the correct search string, reducing the user experience.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] Embodiments of the present disclosure provide a word embedding representation learning method, a word embedding representation learning device, a text recall method, a text recall device, a computer-readable storage medium, and an electronic device, which can, to at least a certain extent, make words with similar morphology have similar distances in the vector space, thereby improving the accuracy and integrity of recall.
[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.
[0007] According to one aspect of the embodiments of the present disclosure, a word embedding representation learning method is provided, including: obtaining a text corpus, performing word segmentation processing on the text corpus, and constructing a graph structure based on the obtained word segments and the pronunciation information corresponding to the word segments; taking each node in the graph structure as an initial node, and randomly walking to obtain a node sequence corresponding to the initial node; training a word embedding representation model according to the node sequence to obtain a word embedding lookup table, and determining a word embedding representation corresponding to the text corpus based on the word embedding lookup table.
[0008] According to one aspect of the embodiments of the present disclosure, there is provided a word embedding representation learning device, including: a graph construction module, configured to obtain a text corpus, perform word segmentation processing on the text corpus, and construct a graph structure based on the obtained word segments and the pronunciation information corresponding to the word segments; a sampling module, configured to use each node in the graph structure as an initial node, and perform random walk to obtain a node sequence corresponding to the initial node; a word embedding obtaining module, configured to train a word embedding representation model according to the node sequence to obtain a word embedding lookup table, and determine a word embedding representation corresponding to the text corpus based on the word embedding lookup table.
[0009] According to one aspect of the embodiments of the present disclosure, there is provided a text recall method, including: obtaining a search string, performing word segmentation processing on the search string to obtain search word segments; querying in a word embedding lookup table according to the search word segments to obtain word embeddings corresponding to the search word segments, where the word embedding lookup table is obtained according to the word embedding representation learning method in the above embodiments; obtaining a search vector corresponding to the search string according to the word embeddings corresponding to all the search word segments, and determining a recalled text according to the search vector and a text vector corresponding to a candidate text.
[0010] According to one aspect of the embodiments of the present disclosure, there is provided a text recall device, including: a word segmentation module, configured to obtain a search string and perform word segmentation processing on the search string to obtain search word segments; a word embedding obtaining module, configured to query in a word embedding lookup table according to the search word segments to obtain word embeddings corresponding to the search word segments, where the word embedding lookup table is obtained according to the word embedding representation learning method in the above embodiments; a recall module, configured to obtain a search vector corresponding to the search string according to the word embeddings corresponding to all the search word segments, and determine a recalled text according to the search vector and a text vector corresponding to a candidate text.
[0011] According to one aspect of the embodiments of the present disclosure, there is provided a computer program product or a computer program, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various alternative implementation manners.
[0012] According to one aspect of the embodiments of the present disclosure, there is provided an electronic device, including: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the methods provided in the above various alternative implementation manners.
[0013] In the technical solutions provided by some embodiments of the present disclosure, by performing word segmentation on a text corpus and constructing a graph structure based on the obtained word segments and the pronunciation information corresponding to the word segments, then performing random walks on the nodes in the graph structure to obtain multiple node sequences, and finally training a word embedding representation model based on the multiple node sequences to obtain a word embedding lookup table, and determining the word embedding representation corresponding to the text corpus based on the word embedding lookup table. On the one hand, the technical solution of the present disclosure can train a word embedding representation model based on a graph structure and introduce pronunciation information into the graph structure, improving the performance of the word embedding representation model, making characters that are similar in form have similar vector representations in the word embedding space, and alleviating the out-of-vocabulary (OOV) problem; on the other hand, it can accurately obtain the word embedding representation corresponding to the text and improve the quality of the recalled text.
[0014] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:
[0016] Figure 1 A schematic diagram showing an exemplary system architecture to which the technical solutions of the embodiments of the present disclosure can be applied;
[0017] Figure 2 A flowchart schematically showing a word embedding representation learning method according to an embodiment of the present disclosure;
[0018] Figure 3 A schematic diagram showing the structure of a graph structure according to an embodiment of the present disclosure;
[0019] Figure 4 A schematic diagram showing the structure of a random walk sampling according to an embodiment of the present disclosure;
[0020] Figure 5 A flowchart schematically showing a text recall method according to an embodiment of the present disclosure;
[0021] Figure 6 A flowchart schematically showing the process of obtaining recalled text according to an embodiment of the present disclosure;
[0022] Figures 7A - 7BSchematically shows a schematic diagram of an enterprise search interface according to an embodiment of the present disclosure;
[0023] Figure 8 Schematically shows a block diagram of a word embedding representation learning device according to an embodiment of the present disclosure;
[0024] Figure 9 Schematically shows a block diagram of a text recall device according to an embodiment of the present disclosure;
[0025] Figure 10 Shows a schematic diagram of the structure of a computer system suitable for implementing the word embedding representation learning device and the text recall device of the embodiments of the present disclosure. Detailed implementation manners
[0026] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0027] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0028] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0029] The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0030] Figure 1 Shows a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present disclosure can be applied.
[0031] As Figure 1As shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. Among them, the terminal device 101 may be a terminal device with a display screen such as a mobile phone, a portable computer, a tablet computer, a desktop computer, etc.; the network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as a wired communication link, a wireless communication link, etc. In the embodiments of the present disclosure, the network 102 between the terminal device 101 and the server 103 may be a wireless communication link, specifically a mobile network.
[0032] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in [description] are merely illustrative. According to the implementation requirements, there may be any number of terminals, networks, and servers. For example, the server 103 may be a server cluster composed of multiple servers, etc., and can be used to store information related to search string processing.
[0033] In one embodiment of the present disclosure, a user inputs a search string through an input device built in or external to the terminal device 101. The search string input by the user can be sent to the server 103 through the network 102. After receiving the search string, the server 103 can first perform word segmentation processing on it to obtain search word segments, and then look up in the word embedding lookup table obtained by pre-training a word embedding representation model according to the search word segments to obtain a search vector corresponding to the search string. Finally, calculate the similarity between the search vector corresponding to the search string and the text vector of the candidate text, and perform text recall according to the similarity to obtain the recalled text corresponding to the search string. When obtaining the word embedding lookup table, first, a text corpus can be obtained, and word segmentation processing is performed on the text corpus to obtain word segments. After obtaining the word segments, a graph structure can be constructed according to the word segments and the pronunciation information corresponding to the word segments. For example, when the text corpus is a Chinese text, the pronunciation information is the pinyin of each character that makes up the word segment, and the pinyin includes the standard pinyin of the character and the pinyin similar to the standard pinyin. Then, each node in the graph structure is used as an initial node, and a node sequence corresponding to each initial node is obtained through a random walk method. Finally, the word embedding representation model is trained according to the node sequence. After the training is completed, the embedding matrix corresponding to the hidden layer in the word embedding representation model can be obtained as the word embedding lookup table, and the word embedding representation corresponding to the text corpus and the word embedding representation corresponding to the search word segments are determined based on the word embedding lookup table. Further, in order to improve the accuracy of word embedding and alleviate the out-of-vocabulary (OOV) problem, a graph structure can be constructed according to common characters and the word library in the business scenario, and the word embedding representation model is trained according to the node sequence determined based on the graph structure to obtain the word embedding lookup table. Furthermore, after obtaining the search string, the word embedding lookup table corresponding to the business scenario can be selected according to the business scenario corresponding to the search string, and the word embedding representation corresponding to the search string can be obtained.
[0034] It should be noted that the word embedding representation learning method and the text recall method provided by the embodiments of the present disclosure are generally executed by the server. Correspondingly, the word embedding representation learning device and the text recall method device are generally set in the server. However, in other embodiments of the present disclosure, the word embedding representation learning method and the text recall method provided by the embodiments of the present disclosure can also be executed by the terminal device.
[0035] In advanced tasks of natural language processing, the use of machine learning methods requires converting words into mathematical representations and then performing calculations using these representations to complete tasks at the semantic level. In statistical learning models, using word embeddings to complete natural language processing tasks is a key technology for natural language processing tasks. In related technologies, common word embedding training methods are mainly divided into two categories: static representation and dynamic representation. Static representation includes word embedding through the bag-of-words model, topic model, classical language model, and optimized language model. Dynamic representation includes word embedding through models such as ELMo (Embeddings from Language Models), GPT, and BERT (Bidirectional Encoder Representation from Transformers). Among them, the bag-of-words model mainly includes discrete representation methods such as one-hot encoding, TF-IDF, and TextRank. However, since the bag-of-words model ignores elements such as the grammar and word order of the document, regarding the document merely as a set of several unordered words and each word being independent, there are problems such as the curse of dimensionality, and there is no correlation relationship between word vectors, resulting in a semantic gap. The topic model mainly includes models based on matrix factorization such as LSA, LDA, and Glove, but there is a problem of large computational complexity. The classical language model mainly includes classical language models such as NPLM and C&W, where word vectors are by-products, but there are problems such as high computational cost and difficulty in engineering implementation. The optimized language model mainly includes targeted optimization models such as word2vec and FastText, but there is a problem of being unable to solve the problem of polysemy. ELMo is a language model for bidirectional semantic feature extraction based on a two-layer bidirectional LSTM, and the main problem is that the feature extraction ability of LSTM is limited and the feature fusion ability of bidirectional splicing is weak. GPT is a unidirectional language model based on the result of the Transformer decoder, with the problem of unidirectional semantics. BERT is a bidirectional language model based on the Transformer encoder structure, with problems such as high training cost and high sample size requirements.
[0036] In view of the problems existing in the related technologies, embodiments of the present disclosure provide a word embedding representation learning method and a text recall method. The word embedding representation learning method and the text recall method are implemented based on machine learning. Machine learning belongs to a type of artificial intelligence. Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0037] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0038] Computer Vision Technology (CV): Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking, and measurement on targets, and further perform graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0039] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0040] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0041] The solution provided by the embodiments of the present disclosure relates to the natural language processing technology of artificial intelligence and can be applied to the information search field. It will be specifically described through the following embodiments:
[0042] Figure 2 Schematically shows a flowchart of a word embedding representation learning method according to an embodiment of the present disclosure. The word embedding representation learning method can be executed by a server, and the server can be Figure 1 the server 103 shown in Figure 2 As shown, the word embedding representation learning method at least includes steps S210 to S230, which are introduced in detail as follows:
[0043] In step S210, a text corpus is obtained, the text corpus is segmented, and a graph structure is constructed based on the obtained segments and the pronunciation information corresponding to the segments.
[0044] In an embodiment of the present disclosure, different countries and regions have different languages, such as Chinese, English, French, German, etc. Although the language types are different, the idea of converting various languages into word vectors is basically the same. In the embodiments of the present disclosure, first, a large amount of text corpus can be obtained. The text corpus can be a corpus related to a specific business scenario, such as a corpus related to enterprise business information, specifically the business registration names of various enterprises, etc.; it can also be a corpus covering multiple related business scenarios, such as a corpus related to insurance and medical examination reports, etc.; of course, it can also be some other corpus, which can be adaptively adjusted according to different needs.
[0045] After obtaining the text corpus, the text corpus can be preprocessed. Specifically, the text corpus can be segmented. For text corpora of different language types, the segmentation methods are slightly different. For example, when the text corpus is a Chinese text, the segmentation can be performed using a dictionary-based segmentation method, a statistics-based segmentation method, or a deep learning-based segmentation method. Specifically, the dictionary-based segmentation method can include the forward maximum matching method, the backward maximum matching method, and the bidirectional maximum matching method. The statistics-based segmentation method is to use a statistical machine learning model to learn the rules of word segmentation (referred to as training) on the premise of a large number of already segmented texts, so as to achieve the segmentation of unknown texts. The main statistical models include: N-gram model, Hidden Markov Model (HMM), Maximum Entropy Model (ME), Conditional Random Fields (CRF), etc. The statistics-based segmentation methods include: N-shortest path method, segmentation method based on word n-gram model, Chinese word segmentation method by constructing words from characters, Chinese word segmentation method based on word perceptron algorithm, Chinese word segmentation method combining character-based generative model and discriminative model. When the text corpus is a non-Chinese text, n-gram segmentation and subword segmentation can be used. The stem of the word in the non-Chinese text can be obtained through n-gram segmentation, and the subwords of the word can be obtained through subword segmentation.
[0046] In one embodiment of the present disclosure, taking information search and recall as an example, when a user inputs a search string, a spelling error or a typo may occur due to negligence, or a recognition error may occur when performing optical character recognition (OCR) on the text. If the search and recall system performs retrieval and recall strictly according to the acquired search string, the recall information may not be obtained or the recall information may be wrong or missing. Therefore, in order to improve the quality of information recall and further improve the user experience, morphologically similar words may have a similar distance in the vector space to ensure that when searching and recalling a large amount of candidate information according to the search string, not only the information containing the search string can be recalled, but also the information containing words morphologically similar to the search string can be recalled. Morphology generally refers to the written form of words and is one of the main elements of written language. Morphologically similar words usually have the same or similar pronunciation information. Taking Chinese characters as an example, the pronunciation information of Teng, Teng, and 縢 are all teng, and the pronunciation information of Xun, Xun, and Xun are all xun. Even non-Chinese characters can have morphologically similar words, such as angel and angle, affect and effect, quite and quiet, etc., and these morphologically similar words also have similar or identical pronunciation information. In order to facilitate the association of morphologically similar words and make morphologically similar words have similar word embedding representations, a graph can be constructed based on morphologically similar words, and vector conversion can be performed based on the graph to obtain the word embedding of the word.
[0047] In one embodiment of the present disclosure, when constructing a graph structure, the segmented words and pronunciation information corresponding to the segmented words obtained by segmenting the text corpus can be used for construction. Specifically, the segmented words and pronunciation information are used as nodes, and the relationship between the segmented words and the relationship between the segmented words and the pronunciation information are used as edges, and then a graph structure is formed according to the nodes and edges. The pronunciation information corresponding to the segmented words is related to the type of the text corpus. When the text corpus is Chinese text, the pronunciation information can be pinyin; when the text corpus is non-Chinese text, the pronunciation information can be phonetic symbols.
[0048] In order to make the technical solution of the present disclosure clearer, a specific description is given below using Chinese text corpus as an example.
[0049] In one embodiment of the present disclosure, a graph structure can be constructed based on the segmentation obtained by the Chinese text corpus through segmentation processing and the pinyin corresponding to each word in the segmentation. Specifically, the segmentation and pinyin corresponding to the Chinese text are nodes, and the relationship between the segmentation, the individual words in the segmentation and the pinyin corresponding to the individual words is an edge, and an undirected acyclic graph is constructed according to the nodes and edges. It is worth noting that the pinyin corresponding to the individual words includes the standard pinyin corresponding to the individual words and the non-standard pinyin, and the non-standard pinyin is similar to the standard pinyin, which is mainly caused by different living areas and different accents, such as the confusion of l and n, and the standard pinyin of milk should be niu nai, but due to non-standard pronunciation, some people will spell it as liu lai, etc., and introduce nodes and edges with different characters with the same pinyin and similar pinyin in the graph structure, which expands the amount of information in the graph structure, and guarantees to the greatest extent that the words with similar morphology are associated, and have a similar distance in the vector space.
[0050] Figure 3 A schematic diagram showing the structure of the graph is shown in FIG. Figure 3 As shown in the figure, the text corpus is segmented to obtain the segmented words "Tencent Cloud", "Tencent", "Tencent", "Yixun", "Xun Teng", "Tengfei" and so on. The individual words in these segmented words are all morphologically similar words. According to the segmented words, the individual words in the segmented words, and the pinyin of the individual words, we can form Figure 3 The graph structure shown. Figure 3 Although the connection relationships between characters with similar pinyin are not shown, it should be understood that the amount of data contained in the graph structure is quite large, and the number of nodes and edges contained therein may be in the tens of millions or even hundreds of millions. Therefore, all connection relationships should be included in the graph structure, such as the edges established on the node relationships between the same characters with different pinyins and the same characters with similar pinyins, which are the key focus.
[0051] In one embodiment of the present disclosure, in the process of constructing a graph structure, weights can be assigned to the edges between each node according to preset rules. The preset rules can be set according to specific business needs. For example, a higher first weight is set for the edge between a segmentation and a single word in the segmentation, a second weight lower than the first weight is set between a single word and standard pinyin, a third weight lower than the second weight is set between a single word and non-standard pinyin, or the weight is determined according to the edit distance of the word in pinyin or composition, such as taking the inverse of the edit distance as the weight of the edge. Of course, a weight model can also be trained separately to assign weights to the edge. The embodiment of the present disclosure does not specifically limit the specific form of the preset rules.
[0052] By constructing a graph structure according to the word segmentation and the pronunciation information corresponding to the word segmentation, and learning the word embeddings of each node based on the graph structure as the distributed representation of characters, high-quality word embeddings can be obtained, enabling words with similar morphology to have high similarity in the word embedding space. Moreover, due to the addition of pinyin nodes, the OOV problem is greatly alleviated. This is because a large number of relevant words and characters are covered in the graph structure, so that the characters in the search string basically fall within the graph structure and there will be no situation beyond the vocabulary.
[0053] In step S220, taking each node in the graph structure as an initial node, random walks are performed to obtain a node sequence corresponding to the initial node.
[0054] In an embodiment of the present disclosure, after the construction of the graph structure is completed, the word embeddings of each node in the graph structure can be learned based on the graph structure through a machine learning model. When learning the word embeddings of each node, first, each node in the graph structure can be used as an initial node, and a node sequence corresponding to the initial node is obtained through random walks, and then the word embeddings are trained according to a large number of node sequences.
[0055] When obtaining the node sequence through random walks, two parameters can be set first: the first parameter p and the second parameter q. p and q are used to achieve a balance between breadth-first search (BFS) and depth-first search (DFS), and local and macroscopic information is considered. Then, according to the first parameter p and the second parameter q, the random walk probabilities of the current node jumping to the historical node and the future node adjacent to the current node are determined. Finally, the random walk direction is determined according to the random walk probabilities, and the node sequence is determined based on the random walk direction.
[0056] Figure 4 Shows a schematic structural diagram of random walk sampling, as Figure 4 shown, there are nodes t, v, x1, x2, and x3 in the graph structure, as well as edges connected to each node. The current node is v, coming from the edge (t, v). It can be analyzed from the figure that when sampling next, the current node v can jump to nodes t, x1, x2, and x3. The random walk probability corresponding to each edge is denoted as α pq (t,x). According to the different nodes connected by the edge and different values of p and q, the random walk probability of each edge is also different. The value of α pq (t,x) can be shown as in formula (1):
[0057]
[0058] where d tx represents the shortest direct path from node t to node x; d tx =0 means returning to node t; dtx = 1 indicates that node t is directly connected to node x, but node v was selected in the previous step; d tx = 2 indicates that node t is not directly connected to node x, but node v is directly connected to node x.
[0059] After determining the parameters p and q, the random walk probability corresponding to each edge can be determined. During sampling, it usually does not return to the nodes that have been sampled. Therefore, the first parameter p is usually set to be relatively large, that is, the probability of walking along the edge (v, t) is very small. Further, when setting the values of the first parameter p and the second parameter q, they can be set according to the sampling requirements. For example, if you mainly want to search in the breadth direction, then q can be set to a value greater than 1, and p can be set to a value greater than q. In this way, during sampling, it will sample along the edge (v, x1); if you want to search in the depth direction, then q can be set to a value greater than 0 and less than 1, and p can be set to a value greater than 1. In this way, during sampling, it will sample along the edges (v, x2) and (v, x3).
[0060] In an embodiment of the present disclosure, during sampling, the sampling length L can be set to obtain multiple node sequences with each node as the initial node and having the sampling length. The sampling length L can be set, for example, as 2 ≤ L ≤ 5. Of course, it can also be set to other numerical ranges according to actual needs.
[0061] In step S230, the word embedding representation model is trained according to the node sequence to obtain a word embedding lookup table, and the word embedding representation corresponding to the text corpus is determined based on the word embedding lookup table.
[0062] In one embodiment of the present disclosure, after obtaining the node sequences with each node in the graph structure as the initial node, the word embedding representation model can be trained according to the node sequences to obtain a stable word embedding representation model and a word embedding lookup table. The word embedding representation model adopted in the present disclosure can specifically be a Node2vec model, etc. The Node2vec model is a model used to generate node vectors in the graph structure. The input is the graph structure generated in step S220, and the output is the vector of each node, that is, the word embedding of the word corresponding to each node. The structure of the Node2vec model includes a Skip-gram model. The Skip-gram model is a type of word2vec model. After obtaining the node sequences, the Skip-gram model can be used to process each node sequence to obtain a prediction result. When training the word embedding representation model according to the node sequences, specifically, the node sequences can be input into the word embedding representation model to obtain the prediction information output by the word embedding representation model, and then the loss function can be determined based on the prediction information and the labeled information corresponding to the node sequences. Finally, the parameters of the word embedding representation model are optimized based on the loss function. When the value of the loss function reaches the minimum or after completing the preset number of training times, the training is considered completed. The Skip-gram model predicts the context words of the target word by inputting the target word, maximizing the probability of the word appearing, that is, the probability of node co-occurrence, where the target word is the word corresponding to any node in the node sequence.
[0063] The Skip-gram model includes an input layer, a hidden layer, and an output layer. Each word in the node sequence is input into the model through the input layer. There is a weight matrix between the input layer and the hidden layer. The value obtained by the hidden layer is the result of the weight matrix acting on the input word. At the same time, there is also a weight matrix from the hidden layer to the output layer. Each value of the output layer vector is the result of the vector of the hidden layer dot-multiplying each column of the weight matrix. Finally, the output layer vector is normalized to obtain the prediction probability of each word, that is, the probability of each word in the vocabulary becoming the context of the target word. The word with the highest probability is the predicted word, that is, the predicted word is the word with the highest co-occurrence probability with the input target word and is the most likely to form a sentence.
[0064] As can be seen from the above process analysis, the key to obtaining word embeddings lies in obtaining the weight matrix from the input layer to the hidden layer. Through the action of this weight matrix, word embeddings can be obtained. After training, the size of this weight matrix is N×M, where N is the vocabulary size and M is the word embedding length. Since each word is assigned a unique number when constructing the vocabulary based on the word segmentation nodes in the graph structure, for example, the words are numbered from 0 to N in sequence, then after obtaining the weight matrix, the vector corresponding to the i-th row in the weight matrix can be found according to the number of the word in the vocabulary to obtain the word embedding corresponding to the word, that is, the i-th row vector in the weight matrix is the word embedding of the i-th word in the vocabulary. Correspondingly, after obtaining the weight matrix, that is, the word embedding lookup table, the numbers of the word segments can be determined according to the word segments corresponding to the text corpus and the vocabulary, and the word embeddings corresponding to each word segment can be obtained in the word embedding lookup table. Furthermore, based on the word embeddings of all the word segments in the text corpus, the word embedding corresponding to the text corpus can be obtained.
[0065] It should be noted that the graph structure in the embodiments of the present disclosure includes words with morphological similarities. Therefore, based on the graph structure for word embedding representation learning, it can make the word embeddings corresponding to words with morphological similarities also have a similar distance in the vector space, making it possible to measure the similarity of homophones. Furthermore, when performing information recall, not only can the information containing the search string be recalled, but also the information containing words that are morphologically similar to the search string can be recalled, avoiding recall errors caused by input errors.
[0066] In an embodiment of the present disclosure, when constructing the graph structure in step S210, the edges between connected nodes have assigned weights, and these weights can act on the loss function when training the model to improve the model performance. The loss function characterizes the degree of difference between the predicted information and the labeled information. When the degree of difference is lower, the loss function is smaller and the model performance is better. Introducing the weights of the edges when calculating the loss function can increase the attention of the model to two nodes with large differences. Furthermore, when adjusting the parameters in the reverse direction, the nodes with differences can be focused on, so that the predicted information output by the optimized model is similar or the same as the labeled information. The loss function can specifically be the cross-entropy loss function, and of course, it can also be other loss functions. The present disclosure does not make specific limitations on this.
[0067] In one embodiment of the present disclosure, during the process of training word embeddings, due to the extremely large amount of text corpus and graph structure data, distributed computing spark is usually used for data processing in engineering. However, there are still the following three problems during the algorithm operation: (1) High graph storage and machine node I / O; (2) Data skew; (3) The data dependency chain is too long in multiple rounds of iteration. For problem (1), when storing the graph structure, nodes and edges need to be stored. If it is stored on multiple machines, the graph structure needs to be divided into multiple subgraphs for storage. When dividing, attention should be paid to the number of nodes and edges. If a machine stores a large number of nodes and cuts edges, when this machine processes the edges, it needs to pull edge information from other machines, which makes the data processing efficiency very low. To solve this problem, a hybrid splitting method can be used for optimization. This hybrid splitting method mainly adopts different splitting strategies according to the node degrees in the graph structure. Specifically, low-degree nodes are split by edges to maintain locality, and high-degree nodes are split by points to reduce node backup, so that the entire graph structure reaches a balance in parallelism and storage. For problem (2), in the text corpus, there may be some words with high frequencies and some words with low frequencies. Then the machines processing the words with high frequencies need to spend a lot of time, while the machines processing the words with low frequencies will quickly complete data processing. However, the data processing logic must wait for all machines to complete their own tasks before the next round of processing can be carried out, which makes the data processing efficiency very low. To solve this problem, it can be alleviated through multi-stage aggregation operations and map join, that is, the tasks of processing the words with high frequencies are divided into multiple subtasks, which are simultaneously executed by multiple machines, and then the processing results of multiple machines are integrated together as a task for processing. For problem (3), since the model training process is a multi-round iteration process, the model performance reaches the optimal through continuous forward propagation and reverse parameter adjustment, that is, the same text corpus will be used for multiple repeated model trainings. This may lead to the machines used to execute the model training algorithm crashing as the number of training times increases, resulting in the failure of the training process. Therefore, a reasonable intermediate variable cache or the persistence of important data structures can be adopted to alleviate this problem and make the whole operation smoother. Specifically, the data dependency chain can be directly cut off, and the intermediate results can be cached. When performing the next data processing, start directly from the intermediate results without repeating the previous process.
[0068] The word embedding representation learning method can solve the word embedding representation problem from the perspective of graph computing. In particular, it enables words with similar morphology to have similar distances in the vector space, making it possible to measure the similarity of homophones in Chinese. Moreover, the word embedding representation learning method in the embodiments of the present disclosure only requires less computing resources to learn large-scale word embeddings on tens of millions of nodes and hundreds of millions of edges, and can be completed within minutes, with high performance.
[0069] Based on the word embedding representation learning method, the present disclosure also provides a text recall method. Figure 5 The flowchart of the text recall method is shown, as Figure 5 described, the method at least includes steps S510 - S530, specifically:
[0070] In step S510, obtain the search string, and perform word segmentation on the search string to obtain search word segments.
[0071] In an embodiment of the present disclosure, the user inputs a search string through an input device in the terminal interface. The search string can be in Chinese, English, or other types of strings. In the embodiments of the present disclosure, still taking a Chinese search string as an example for illustration, the search string can be, for example, a person's name, and information about the person's name matching the search string is obtained according to the search string; it can be, for example, an enterprise name, and information about the corresponding enterprise is queried on the industrial and commercial enterprise registration platform according to the search string, and so on.
[0072] In an embodiment of the present disclosure, before performing search recall according to the search string, it is necessary to preprocess the search string, that is, perform word segmentation on the search string to obtain search word segments. For example, if the search string is the enterprise name "XX Technology Co., Ltd.", the search word segments obtained through word segmentation are "XX Technology Co., Ltd.". After obtaining the search word segments, the word embedding corresponding to the search word segments can be determined based on the word embedding lookup table.
[0073] In step S520, query in the word embedding lookup table according to the search word segments to obtain the word embedding representation corresponding to the search word segments. The word embedding lookup table is the word embedding lookup table obtained according to the word embedding representation learning method in the above embodiments.
[0074] In an embodiment of the present disclosure, when the word embedding lookup table contains embedding vectors of a sufficient number of words, the word embedding corresponding to the search word segments can be obtained from the word embedding lookup table. Specifically, first determine the encoding of the search word segments according to the word list, and then use the encoding of the search word segments as an index to search for the corresponding embedding vector in the word embedding lookup table. This embedding vector is the word embedding representation of the search word segments.
[0075] For the word embedding lookup table to contain embedding vectors of a sufficient number of words, on the one hand, it is necessary to collect corpora that can cover almost all business scenarios, and on the other hand, it is necessary to construct a huge graph structure based on the corpora, which poses great challenges to the storage and processing efficiency of the machine. Therefore, in order to further improve the data processing efficiency and avoid the OOV problem, a graph structure can be constructed based on the text corpora of different business scenarios and word embedding training can be carried out to obtain word embedding lookup tables corresponding to different business scenarios. After obtaining the search string, the business scenario corresponding to the search string can be determined, and the corresponding target word embedding lookup table can be determined in the word embedding lookup tables corresponding to different business scenarios according to the business scenario. Then, the word embedding corresponding to the search token can be retrieved from the target word embedding lookup table according to the search tokenization. This can not only improve the efficiency and quality of model training, but also improve the efficiency of converting the search string into a vector.
[0076] In step S530, a search vector corresponding to the search string is obtained according to the word embeddings corresponding to all the search tokens, and a recalled text is determined according to the search vector and the text vector corresponding to the candidate text.
[0077] In an embodiment of the present disclosure, after obtaining the word embeddings corresponding to each search token in the search string, the word embeddings corresponding to all the search tokens can be concatenated in order to obtain a search vector corresponding to the search string, and then the search vector is matched with the text vector corresponding to the candidate text to obtain the recalled text.
[0078] In an embodiment of the present disclosure, during text recall, the number of candidate texts is usually multiple. Then, when determining the recalled text according to the search vector and the text vector corresponding to the candidate text, the first similarity between the search vector and the text vectors of each candidate text can be calculated, and the recalled text can be determined according to the first similarity. When the first similarity is greater than or equal to the preset similarity threshold, the candidate text is recalled as the recalled text. When the first similarity is less than the preset similarity threshold, the candidate text is filtered out. The first similarity can be determined by calculating distances such as the cosine distance, Euclidean distance, and Hamming distance between the search vector and the text vector. The higher the first similarity, the more closely the corresponding candidate text matches the search string. Since in the process of word embedding representation learning, the word embeddings corresponding to morphologically similar words have similar distances in the vector space, when determining the recalled text according to the first similarity, not only the text containing the search string can be recalled, but also the text containing words that are morphologically similar to the search string can be recalled, avoiding the situation of incorrect or missing recalled text caused by typos in the search string, etc., and thus improving the user experience.
[0079] In one embodiment of the present disclosure, when the search string and the candidate text only contain morphologically similar words, the recall can be performed in the manner of the above embodiments, such as person name recall, product recall, etc. By obtaining the word embeddings of the search person name and the candidate person name, and the word embeddings of the search product name and the candidate product name according to the method in the embodiment of the present disclosure, and then calculating the similarity between the word embeddings of the search person name and the candidate person name for person name recall, or calculating the similarity between the word embeddings of the search product name and the candidate product name for product recall. However, when the search string and the candidate text contain multiple fields with different attributes, recall needs to be performed separately according to the fields with different attributes. For example, if the search string contains not only morphologically similar words but also semantically similar words, then recall cannot be performed simply by obtaining word embeddings and calculating similarities according to the method in the embodiment of the present disclosure. Figure 6 shows a schematic flow diagram for obtaining recall text, as Figure 6 shown. In step S601, an inverted index is performed on the candidate text and the text vector corresponding to the candidate text, and the second similarity between the search vector and each text vector is determined, and initial recall is performed according to the second similarity; in step S602, the third similarity between the search string and the vectors corresponding to the fields with the same attributes in the candidate text obtained by the initial recall is obtained, and re-recall is performed on the candidate text obtained by the initial recall according to the third similarity to obtain the recall text. Among them, the calculation methods of the second similarity and the third similarity may be the same as or different from the calculation method of the first similarity, and the embodiments of the present disclosure do not make specific limitations on this.
[0080] Taking enterprise search as an example, Figures 7A - 7B shows a schematic interface diagram of enterprise search, as Figure 7AAs shown, the user enters the name of the enterprise to be searched in the display interface of the terminal. For example, the user enters "Tencent Technology (Beijing) Co., Ltd.". After receiving this search enterprise name, first, the enterprise name can be divided into four segments by the sequence labeling model: Tencent, Technology, Beijing, Co., Ltd. Among them, Tencent is the enterprise name, Technology is the industry attribute of the enterprise, Beijing is the geographical location attribute of the enterprise, and Co., Ltd. is the basic attribute of the enterprise. Then, when querying the enterprise information corresponding to this enterprise name, it is necessary to query and recall from these four attributes. Among these four fields, only the enterprise name involves the problem of morphological similarity. For example, the user actually wants to search for "Tencent Technology (Beijing) Co., Ltd.", but enters "Tengxun Technology (Beijing) Co., Ltd." when inputting. As for the industry attribute and basic attribute, they mainly involve semantic problems. For example, technology and technique are similar in semantics, and Co., Ltd. and Limited Liability Company are similar in semantics. Then, different vector conversion methods can be used to encode the fields with different attributes in the enterprise name. Among them, for the enterprise name, word embedding conversion can be used to obtain the word embedding lookup table related to the enterprise search business scenario by the word embedding representation learning method in this embodiment of the present disclosure, and then determine its corresponding word embedding in the word embedding lookup table according to the encoding of the enterprise name in the word table. After determining the vectors corresponding to each field, the search vector corresponding to the search enterprise name can be obtained. When querying in the enterprise information query platform, the search vector corresponding to the search enterprise name can be matched with the text vectors corresponding to the candidate enterprise names stored in the database, and the candidate enterprise names obtained by the matching are returned to the terminal for the user to click to view the enterprise details.
[0081] Among them, the method for obtaining the text vectors corresponding to the candidate enterprise names is the same as that of the search vectors, which will not be elaborated here. When matching the search vector and the text vectors corresponding to the candidate enterprise names, first match in the full space. Specifically, first perform an inverted index on the candidate enterprise names and the corresponding text vectors, and then determine the similarity between the search vector corresponding to the search enterprise name and the text vectors corresponding to each candidate enterprise name, and recall the candidate enterprise names with a similarity greater than the preset threshold to achieve the initial recall. After the initial recall, the similarity between the vectors corresponding to the fields with the same attributes in the search enterprise name and the initially recalled candidate enterprise names can be determined, and then the similarities corresponding to each attribute are sorted. According to the preset similarity threshold, the candidate enterprise names corresponding to each attribute are obtained, and the common candidate enterprise names among them are recalled as the search results and fed back to the user, as Figure 7B shown.
[0082] Based on the word embedding representation learning method and text recall method in the present disclosure, texts containing the search string can be recalled, and texts containing characters that are morphologically similar to the search string can also be recalled, improving the recall quantity and quality, avoiding inaccurate and incomplete recall information caused by input errors or recognition errors of the search string, and further improving the user experience.
[0083] The following introduces the apparatus embodiments of the present disclosure, which can be used to execute the word embedding representation learning method and text recall method in the above embodiments of the present disclosure. For details not disclosed in the apparatus embodiments of the present disclosure, please refer to the embodiments of the word embedding representation learning method and text recall method above of the present disclosure.
[0084] Figure 8 The block diagram of a word embedding representation learning apparatus according to an embodiment of the present disclosure is schematically shown.
[0085] Referring to Figure 8 As shown, a word embedding representation learning apparatus 800 according to an embodiment of the present disclosure includes: a graph construction module 801, a sampling module 802, and a word embedding acquisition module 803.
[0086] Among them, the graph construction module 801 is configured to obtain a text corpus, perform word segmentation processing on the text corpus, and construct a graph structure based on the obtained word segmentation and the pronunciation information corresponding to the word segmentation; the sampling module 802 is configured to use each node in the graph structure as an initial node and randomly walk to obtain a node sequence corresponding to the initial node; the word embedding acquisition module 803 is configured to train a word embedding representation model according to the node sequence to obtain a word embedding lookup table, and determine a word embedding representation corresponding to the text corpus based on the word embedding lookup table.
[0087] In an embodiment of the present disclosure, the text corpus is a Chinese text, and the pronunciation information is the pinyin corresponding to each character in each word segmentation obtained by performing word segmentation processing on the Chinese text; the graph construction module 801 is configured to: use the word segmentation corresponding to the Chinese text and the pinyin as nodes, use the relationship between the word segmentation, single characters in the word segmentation, and the pinyin corresponding to the single characters as edges, and construct an undirected acyclic graph according to the nodes and the edges.
[0088] In an embodiment of the present disclosure, the graph construction module 801 is further configured to: when constructing the undirected acyclic graph, set weights for each of the edges according to a preset rule.
[0089] In an embodiment of the present disclosure, the edges include edges established on the node relationships where the pinyins are the same but the characters are different and where the pinyins are similar and the characters are the same.
[0090] In one embodiment of the present disclosure, the sampling module 802 is configured to: obtain a preset first parameter and a second parameter, determine the transition probabilities of the current node jumping to the historical node and the current node jumping to the future node according to the current node, the historical node and the future node adjacent to the current node, the first parameter and the second parameter; determine a transition direction according to the transition probabilities, and determine the node sequence based on the transition direction.
[0091] In one embodiment of the present disclosure, the word embedding obtaining module 803 includes: a prediction information obtaining unit, configured to input the node sequence into the word embedding representation model to obtain prediction information; a loss function determining unit, configured to determine a loss function according to the prediction information and the labeled information corresponding to the node sequence; a parameter optimization unit, configured to optimize the parameters of the word embedding representation model based on the loss function so that the value of the loss function reaches the minimum, and use the embedding matrix corresponding to the hidden layer in the trained word embedding representation model as the word embedding lookup table.
[0092] In one embodiment of the present disclosure, the word embedding obtaining module 803 is configured to: obtain a word list constructed based on the graph structure, and obtain the encoding corresponding to the word segmentation in the text corpus according to the word list; determine the word embedding corresponding to the word segmentation in the word embedding lookup table according to the encoding; determine the word embedding representation corresponding to the text corpus according to the word embeddings corresponding to all the word segmentations.
[0093] Figure 9 A block diagram of a text recall device according to an embodiment of the present disclosure is schematically shown.
[0094] Referring to Figure 9 As shown, a text recall device 900 according to an embodiment of the present disclosure includes: a word segmentation module 901, a word embedding obtaining module 902, and a recall module 903.
[0095] Among them, the word segmentation module 901 is configured to obtain a search string and perform word segmentation processing on the search string to obtain search word segments; the word embedding obtaining module 902 is configured to query in the word embedding lookup table according to the search word segments to obtain the word embeddings corresponding to the search word segments, and the word embedding lookup table is the word embedding lookup table obtained according to the word embedding representation learning method in the above embodiment; the recall module 903 is configured to obtain a search vector corresponding to the search string according to the word embeddings corresponding to all the search word segments, and determine a recalled text according to the search vector and the text vector corresponding to the candidate text.
[0096] In one embodiment of the present disclosure, the word embedding acquisition module 902 is configured to: determine the service scenario corresponding to the search string, and determine the target word embedding lookup table according to the service scenario; query in the target word embedding lookup table according to the search word segmentation to obtain the word embedding corresponding to the search word segmentation.
[0097] In one embodiment of the present disclosure, the number of candidate texts is multiple; the recall module 903 includes: a recall unit, configured to obtain a first similarity between the search vector and the text vectors corresponding to the candidate texts, and determine the recalled text according to the first similarity.
[0098] In one embodiment of the present disclosure, the search string and the candidate texts include fields with multiple different attributes; the recall unit is configured to: perform an inverted index according to the candidate text and the text vector, and determine a second similarity between the search vector and each text vector, and perform an initial recall according to the second similarity; obtain a third similarity between the search string and the vectors corresponding to the fields with the same attributes in the candidate texts obtained by the initial recall, and perform a re-recall in the results of the initial recall to obtain the recalled text.
[0099] Figure 10 The structure diagram of the computer system of the electronic device suitable for implementing the embodiments of the present disclosure is shown.
[0100] It should be noted that Figure 10 The computer system 1000 of the shown electronic device is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present disclosure.
[0101] As Figure 10 shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003, and implement the search string processing method described in the above embodiments. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0102] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read therefrom is installed into the storage section 1008 as needed.
[0103] Specifically, according to an embodiment of the present disclosure, the processes described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by a central processing unit (CPU) 1001, various functions defined in the system of the present disclosure are executed.
[0104] It should be noted that the computer-readable medium shown in the embodiments of the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0106] The units involved in the embodiments described in this disclosure can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.
[0107] As another aspect, the present disclosure also provides a computer-readable medium, which can be included in the word embedding representation learning device and the text recall device described in the above embodiments; or can exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiments.
[0108] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0109] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0110] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure.
[0111] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for learning word embedding representation, characterized in that Including: Obtain a text corpus and perform word segmentation processing on the text corpus; The text corpus is a Chinese text, and the pronunciation information is the pinyin corresponding to each Chinese character in each word obtained by word segmentation of the Chinese text; Using the words obtained by word segmentation corresponding to the Chinese text and the pinyin as nodes, and using the relationships among the words, single characters in the words, and the pinyin corresponding to the single characters as edges, construct an undirected acyclic graph according to the nodes and the edges; Divide each node in the undirected acyclic graph into high-degree nodes and low-degree nodes according to the number of edges corresponding to each node; Use an edge splitting method to split the low-degree nodes in the undirected acyclic graph, and use a point splitting method to split the high-degree nodes in the undirected acyclic graph to obtain multiple subgraphs of the undirected acyclic graph, and store the multiple subgraphs separately; Using each node in the undirected acyclic graph as an initial node, perform random walks to obtain a node sequence corresponding to the initial node; Train a word embedding representation model according to the node sequence to obtain a word embedding lookup table; In each round of training of the word embedding representation model, the processing result of the high-frequency words in the text corpus is obtained by integrating the processing results of multiple subtasks. The multiple subtasks are obtained by dividing the processing task of the high-frequency words, and each subtask is executed by different machines simultaneously; Obtain a vocabulary constructed based on the undirected acyclic graph, and obtain the encoding corresponding to the words obtained by word segmentation in the text corpus according to the vocabulary; Determine the word embedding corresponding to the word obtained by word segmentation in the word embedding lookup table according to the encoding; Determine the word embedding representation corresponding to the text corpus according to the word embeddings corresponding to all the words obtained by word segmentation.
2. The method according to claim 1, characterized in that, The method further includes: When constructing the undirected acyclic graph, set weights for each of the edges according to a preset rule.
3. The method according to claim 1, characterized in that The edges include edges established on the node relationships where the pinyins are the same but the characters are different and where the pinyins are similar and the characters are the same.
4. The method according to claim 1, wherein The step of using each node in the graph structure as an initial node and performing random walks to obtain a node sequence corresponding to the initial node includes: Obtain a preset first parameter and second parameter, and determine the random walk probabilities of the current node jumping to the historical node and the current node jumping to the future node according to the current node, the historical node adjacent to the current node, the future node, the first parameter, and the second parameter; Determine the random walk direction according to the random walk probabilities, and determine the node sequence based on the random walk direction.
5. The method according to claim 1, characterized in that The step of training a word embedding representation model according to the node sequence to obtain a word embedding lookup table includes: Input the node sequence into the word embedding representation model to obtain prediction information; Determine a loss function according to the prediction information and the labeled information corresponding to the node sequence; Optimize the parameters of the word embedding representation model based on the loss function so that the value of the loss function reaches the minimum, and use the embedding matrix corresponding to the hidden layer in the trained word embedding representation model as the word embedding lookup table.
6. A text recall method, characterized in that, Including: Obtain a search string, perform word segmentation processing on the search string to obtain search words obtained by word segmentation; Query in the word embedding lookup table according to the search word segmentation to obtain the word embedding corresponding to the search word segmentation, where the word embedding lookup table is the word embedding lookup table obtained according to the word embedding representation learning method described in any one of claims 1-5; Obtain a search vector corresponding to the search string according to the word embeddings corresponding to all the search word segmentations, and determine the recalled text according to the search vector and the text vector corresponding to the candidate text.
7. The method according to claim 6, characterized in that, The querying in the word embedding lookup table according to the search word segmentation to obtain the word embedding corresponding to the search word segmentation includes: Determine the business scenario corresponding to the search string, and determine the target word embedding lookup table according to the business scenario; Query in the target word embedding lookup table according to the search word segmentation to obtain the word embedding corresponding to the search word segmentation.
8. The method according to claim 6, wherein The number of the candidate texts is multiple; The determining the recalled text according to the search vector and the text vector corresponding to the candidate text includes: Obtain a first similarity between the search vector and the text vectors corresponding to the candidate texts, and determine the recalled text according to the first similarity.
9. The method according to claim 8, wherein The search string and the candidate text include fields with multiple different attributes; The obtaining the first similarity between the search vector and the text vectors corresponding to the candidate texts, and determining the recalled text according to the first similarity includes: Perform an inverted index according to the candidate text and the text vector, and determine a second similarity between the search vector and each text vector, and perform an initial recall according to the second similarity; Obtain a third similarity between the search string and the vectors corresponding to the fields with the same attributes in the candidate texts obtained by the initial recall, and perform a re-recall in the candidate texts obtained by the initial recall to obtain the recalled text.
10. A word embedding representation learning device, characterized in that, including: A graph construction module, configured to obtain a text corpus, perform word segmentation processing on the text corpus, and construct a graph structure based on the obtained word segmentation and the pronunciation information corresponding to the word segmentation; the text corpus is a Chinese text, and the pronunciation information is the pinyin corresponding to each Chinese character in each word segmentation obtained by performing word segmentation processing on the Chinese text; Use the word segmentation corresponding to the Chinese text and the pinyin as nodes, use the relationships between the word segmentation, the single characters in the word segmentation, and the pinyin corresponding to the single characters as edges, and construct an undirected acyclic graph according to the nodes and the edges; Divide each node in the undirected acyclic graph into a high-degree node and a low-degree node according to the number of edges corresponding to each node; Use an edge splitting method to split the low-degree nodes in the undirected acyclic graph, and use a node splitting method to split the high-degree nodes in the undirected acyclic graph to obtain multiple subgraphs of the undirected acyclic graph, and store the multiple subgraphs separately; A sampling module, configured to use each node in the undirected acyclic graph as an initial node, and randomly walk to obtain a node sequence corresponding to the initial node; A word embedding obtaining module, configured to train a word embedding representation model according to the node sequence to obtain a word embedding lookup table; In each round of training of the word embedding representation model, the processing result of the high-frequency words in the text corpus is obtained by integrating the processing results of multiple subtasks, and the multiple subtasks are obtained by dividing the processing task of the high-frequency words, and each subtask is executed simultaneously by different machines; Obtain a vocabulary constructed based on the directed acyclic graph, and obtain the encoding corresponding to the word segmentation in the text corpus according to the vocabulary; Determine the word embedding corresponding to the word segmentation in the word embedding lookup table according to the encoding; Determine the word embedding representation corresponding to the text corpus according to the word embeddings corresponding to all the word segmentations.
11. A text recall device, characterized in that, It includes: A word segmentation module, configured to obtain a search string, perform word segmentation processing on the search string to obtain search word segments; A word embedding acquisition module, configured to query in a word embedding lookup table according to the search word segments to obtain the word embedding corresponding to the search word segments, and the word embedding lookup table is the word embedding lookup table obtained according to the word embedding representation learning method described in any one of claims 1-5; A recall module, configured to obtain a search vector corresponding to the search string according to the word embeddings corresponding to all the search word segments, and determine a recalled text according to the search vector and the text vector corresponding to the candidate text.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the word embedding representation learning method described in any one of claims 1 to 5 and the text recall method described in any one of claims 6 to 9.
13. An electronic device, characterized in that, It includes: One or more processors; A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the word embedding representation learning method described in any one of claims 1 to 5 and the text recall method described in any one of claims 6 to 9.
14. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are adapted to be loaded and executed by a processor to implement the word embedding representation learning method described in any one of claims 1 to 5 and the text recall method described in any one of claims 6 to 9.