Enhanced Method for Proximity Information Retrieval in a Medical Q&A System
Calculating sentence alignment scores through word embedding model and improved Needman-Wonsch algorithm, solving the accuracy of sentence similarity assessment in the biomedical Q&A system, and improving the relevance ranking and document representativeness of information retrieval.
Patent Information
- Application Number
- CN202080028799.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-24
- Filing Date
- 2020-03-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-03-27
AI Technical Summary
In the existing biomedical question and answer system, information retrieval methods based on bag-of-word models are difficult to effectively distinguish sentences with similar semantics but different contexts, resulting in inaccurate correlation rankings.
The word embedding model is used to generate the vector set of sentences, and the sentence alignment score is calculated through the similarity matrix and the improved Needman-Wonsch algorithm to enhance the correlation ranking of information retrieval.
Improve the context similarity evaluation of sentence matching, increase the representativeness and accuracy of the retrieved documents, and support the effective processing of downstream medical QA systems.
Smart Images

Figure CN113711204B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Patent Application No. 16 / 421,554, filed on May 24, 2019, with the United States Patent and Trademark Office, which is hereby incorporated by reference in its entirety. Background Art
[0003] A question - answering (QA) system is a system designed to answer questions posed in available information such as images, videos, voice, and natural language. As an example, a medical QA system extracts information from unified biomedical literature and is designed to answer medical - related questions. Continuing with this example, the knowledge base is concentrated on biomedical resources, and questions and answers are expressed in full text given in Mandarin. Parsing a QA question requires several basic capabilities, including information retrieval (IR), reasoning, memory, etc., and the most encountered, uncertain, and crucial step is the IR step. For example, the task is to select the most relevant references from millions of documents in the knowledge base. Usually, the problem is not due to limited biomedical resources, but due to defects associated with ranking and identifying the context - most - relevant resources for questions and answers from millions of available resources.
[0004] It has been generally recognized that one of the full - text search engines, namely Lucene (keyword analyzer "Search better with Apache Lucene and Solr", November 19, 2007), can implement a recommendation system (McCandless, Michael; Hatcher, Eik; Gospodnetic, Otis (2010), Lucene in Action, 2nd Edition. Manning, page 8. ISBN 193398817). The power of the inverted index and the relevance ranking derived from term frequency - inverse document frequency (TF - IDF) (such as Okapi BM25) have been shown to be able to rank documents based on query terms in the bag - of - words (BOW) method, regardless of the inter - relationships between the matches within the document. This BOW method is a feature where search can be efficiently implemented based on the cosine model, and at the same time, documents can be reasonably ranked in many recommendation systems. However, the BOW method also has its own drawbacks. One example of use is when searching for two similar sentences with almost the same BOW but different contextual meanings. It may become increasingly difficult to distinguish between these two similar sentences based only on the relevance score. To solve this problem of implementing a better IR component in a biomedical QA system, the present disclosure provides enhanced proximity search extended from ElasticSearch / Lucene indexes. Summary of the Invention
[0005] According to one aspect of the present disclosure, a method for performing information retrieval using sentence similarity, the method comprising: receiving, by a device, a first sentence including a first set of words; receiving, by the device, a second sentence including a second set of words; generating, by the device and using a word embedding model, a first set of vectors corresponding to the first set of words of the first sentence; generating, by the device and using the word embedding model, a second set of vectors corresponding to the second set of words of the second sentence; generating, by the device, a similarity matrix based on the first set of vectors and the second set of vectors; determining, by the device, an alignment score associated with the first set of vectors and the second set of vectors using the similarity matrix; and sending, by the device, the alignment score to allow information retrieval to be performed based on the similarity between the first sentence and the second sentence.
[0006] According to one aspect of the present disclosure, a device comprising: at least one memory configured to store program code; at least one processor configured to read the program code and operate in accordance with the instructions of the program code, the program code including: receiving code configured to cause the at least one processor to receive a first sentence including a first set of words and a second sentence including a second set of words; generating code configured to cause the at least one processor to use a word embedding model to generate a first set of vectors corresponding to the first set of words of the first sentence, use the word embedding model to generate a second set of vectors corresponding to the second set of words of the second sentence, and generate a similarity matrix based on the first set of vectors and the second set of vectors; determining code configured to cause the at least one processor to use the similarity matrix to determine an alignment score associated with the first set of vectors and the second set of vectors; and sending code configured to cause the at least one processor to send the alignment score to allow information retrieval to be performed based on the similarity between the first sentence and the second sentence.
[0007] According to one aspect of the present disclosure, a non-transitory computer-readable medium stores instructions, the instructions including one or more instructions that, when run by one or more processors of a device, cause the one or more processors to: receive, by the device, a first sentence including a first set of words; receive, by the device, a second sentence including a second set of words; generate, by the device and using a word embedding model, a first set of vectors corresponding to the first set of words of the first sentence; generate, by the device and using the word embedding model, a second set of vectors corresponding to the second set of words of the second sentence; generate, by the device, a similarity matrix based on the first set of vectors and the second set of vectors; determine, by the device, an alignment score associated with the first set of vectors and the second set of vectors using the similarity matrix; and send, by the device, the alignment score to allow information retrieval to be performed based on the similarity between the first sentence and the second sentence. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is a flowchart of an example method described herein;
[0009] Figure 2 is a diagram of an example environment that can implement the systems and / or methods described herein; and
[0010] Figure 3 is Figure 2 a diagram of example components of one or more devices of DETAILED DESCRIPTION
[0011] Recent developments in natural language processing (NLP) and IR technologies have facilitated the analysis of large-scale digital biomedical information (such as images, text, clinical, and genetic data). Advances in biomedical NLP, such as term entity recognition (Kim S, Lu Z, Wilbur WJ, Identifying named entities from PubMed for enriched semantic categories. BMC Bioinformatics journal, February 21, 2015; 16:57); the availability of inverted index search engines, such as Lucene 2 ; and distributed column-based storage systems and analytical databases, such as BigTable (Chang, Fay; Dean, Jeffrey; Ghemawat, Sanjay et al. (published in 2006), "BigTable: A Distributed Storage System for Structured Data", Google) and ElasticSearch (Elasticsearch, https: / / www.elastic.co / ), have enabled biomedical searches and are even more capable of handling billions of documents, supporting high concurrent queries, and returning document relevance related to the query.
[0012] The IR component of the disclosed biomedical QA system is designed on top of ElasticSearch (Elasticsearch) and Lucene indexes, where queries and documents can be analyzed by stemming, performing stop word analysis, and performing synonym expansion. However, an obvious problem is caused by the BOW search strategy of boolean queries in the Lucene index, where the proximity of two matching documents is not easily and directly reflected by the relevance score.
[0013] Due to the complex nature of language and high-volume documents, the full text content may contain similar bags of words, while the contextual differences in the bags of words can be significantly different. On the other hand, in order to understand the meaning of the text, the interpretation of sentences becomes increasingly important, which is particularly crucial in a QA system. The following is an example given in Mandarin:
[0014] 1. When a Class A infectious disease is found in a medical institution, suspected patients shall be isolated and treated separately in a designated place before being diagnosed;
[0015] 2. Where a close contact of a patient, pathogen carrier or suspected patient with Class A infectious diseases within a medical institution is discovered, medical observation shall be conducted in a designated place and other necessary preventive measures shall be taken.
[0016] If the question is to inquire about a better answer by querying "Measures for a suspected patient with Class A infectious diseases discovered within a medical institution", and the two possible answers are "Isolated treatment separately" and "Medical observation", it is difficult to tell which answer is correct based solely on the default score because the search terms are very close and similar in terms of their bag-of-words presentation. Proximity search using span queries can help. However, the terms from the two queries need to contain all the query terms, which is a criterion that is usually not met during the full-text search process. In the present disclosure, a post-processing step is added to analyze the span of the term groups in the returned matching documents and rank the peer documents based on the proximity of the terms.
[0017] The proposed disclosure can be used to enhance the IR component in a biomedical QA system for selecting the most relevant and representative references for use with attention and prediction models. The effectiveness and richness of this IR component are important for the conceptual interpretation of the QA system. Relevance ranking is an important factor in evaluating the importance of references, which helps to understand the question and relate the concept of the question to the correct answer. The present disclosure provides a relevance ranking based on the proximity of the full text of the question to the full text of the answer to retrieve references as close as possible to the QA.
[0018] The full texts of the question and the answer are analyzed through an NLP process, in which tokenization and synonym expansion occur during the Elastic Search indexing time. The full text is tokenized based on bigrams and trigrams, stop-word analysis is applied after tokenization, and then acronyms are expanded. The acronyms and the thesaurus are based on Cilin synonyms (https: / / githu.com / BiLiangLtd / WordSimilarity / blob / master / dat / cilin_ex.txt) and are further expanded by modeling approximately 300,000 documents collected from nearly 100 biomedical literature resources. The titles and full contents of the literature are concatenated to reduce the gap between the terms across the titles and the contents.
[0019] Before adjusting the sentence similarity, word vectorization is an important step in measuring the similarity between words. This disclosure provides word2vec (https: / / code.google.com / archive / p / word2vec / ) implementations of the continuous bag of words and skip-gram architectures for computing the vector representations of words. Thus, the similarity between two words can be measured by cosine similarity (https: / / en.wikipedia.org / wiki / Cosine_similarity):
[0020]
[0021] This value ranges between 0 and 1, where the closer the value is to 1, the more similar the words are indicated to be.
[0022] One way to measure string similarity is to measure its edit distance by counting the minimum number of operations required to convert one string into another. For example, the Levenshtein distance (https: / / en.wikipedia.org / wiki / Levenshtein_distance) measures operations including removing, inserting, or replacing characters in a string. In the case of full text similarity where there may be hundreds of words across multiple sentences and where the sentences being compared may not have good word recognition coverage, edit distance may provide less useful information compared to TF-IDF. This disclosure presents an improved Needleman-Wunsch algorithm that uses dynamic programming to align two sentences, where the score for a replacement uses the cosine score represented by word2vec.
[0023] Adopt the following algorithm pseudocode and modify it from the original implementation (https: / / en.wikipedia.org / wiki / Needleman%E2%80%93Wunsch_algorithm), where the F matrix is a matrix that holds the alignment scores for two lists of words. F(i, j) is the score for the match / mismatch of two words (each word from a different sentence). A and B are the vectors of words from the two sentences to be compared. The "similarity" function takes the vector representations of two words for computing the cosine similarity. The algorithm involves forward and backward passes on the F matrix.
[0024] The forward pass is for computing the F matrix:
[0025]
[0026]
[0027] Once the F matrix is calculated, backtracking is performed to aggregate the alignment starting from the bottom-right cell, and the values in three possible movement directions (top, left, upper-left diagonal) are compared to see which channel gives the best score:
[0028]
[0029]
[0030] The cumulative scores “score” ([0, 1]) and IdentityScore ([0, 1]) can be used to evaluate sentence similarity.
[0031] In this way, the proposed proximity search-based extension of TF-IDF-based correlation scores allows higher-ranked documents to indicate their contextual similarity.
[0032] The similarity between tokens is not limited to exact replacement, but can also be the meaning of their proximity (word2vec method), which makes the full-text interpretation closer to the meaning used to facilitate the QA system.
[0033] The up-weighting of the proximity of matches in terms of extracted item matching allows more diverse documents to be ranked at the top, which increases the representativeness of the retrieved documents for downstream medical QA system processing.
[0034] Figure 1 is a flowchart of an example method 100 according to one aspect of the present disclosure.
[0035] As Figure 1 shown, method 100 may include: receiving, by the device, a first sentence including a first set of words (block 110).
[0036] Further as Figure 1 shown, method 100 may include: receiving, by the device, a second sentence including a second set of words (block 120).
[0037] Further as Figure 1 shown, method 100 may include: generating, by the device and using a word embedding model, a first set of vectors corresponding to the first set of words of the first sentence (block 130).
[0038] Further as Figure 1 shown, method 100 may include: generating, by the device and using a word embedding model, a second set of vectors corresponding to the second set of words of the second sentence (block 140).
[0039] Further as Figure 1 shown, method 100 may include: generating, by the device, a similarity matrix based on the first set of vectors and the second set of vectors (block 150).
[0040] Further, as Figure 1 shown, method 100 may include: determining an alignment score associated with a first vector set and a second vector set using a similarity matrix by a device (block 160).
[0041] Further, as Figure 1 shown, method 100 may include: determining whether the alignment score is the maximum alignment score (block 170).
[0042] Further, as Figure 1 shown, if the alignment score is not the maximum alignment score (block 170 - No), then method 100 may include: determining another alignment score based on different directions of the similarity matrix (block 180).
[0043] Further, as Figure 1 shown, if the alignment score is the maximum alignment score (block 170 - Yes), then method 100 may include: sending the alignment score by the device to allow information retrieval to be performed based on the similarity between the first sentence and the second sentence (block 190).
[0044] In some implementations, Figure 1 one or more process blocks of Figure 2 may be performed by platform 220 as described in connection with Figure 1 . In some implementations,
[0045] one or more process blocks of Figure 1 may be performed by another device or group of devices separate from or including platform 220, such as by user device 210.
[0045] Although Figure 1 shows example blocks of method 100, in some implementations, method 100 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to the blocks depicted in Figure 1 . Additionally or alternatively, two or more blocks of method 400 may be performed in parallel.
[0046] Figure 2 is a diagram of an example environment 200 that may implement the systems and / or methods described herein. As Figure 2 shown, environment 200 may include user device 210, platform 220, and network 230. The devices of environment 200 may be interconnected by a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.
[0047] User device 210 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with platform 220. For example, user device 210 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., a pair of smart glasses or a smartwatch), or a similar device. In some implementations, user device 210 may receive information from platform 220 and / or send information to platform 220.
[0048] Platform 220 includes one or more devices capable of using artificial intelligence (AI) techniques to identify bug bites, as described elsewhere in this document. In some implementations, platform 220 may include a cloud server or a group of cloud servers. In some implementations, platform 220 may be designed as a modular platform such that certain software components can be swapped in or out according to specific needs. Thus, platform 220 can be easily and / or quickly reconfigured for different uses.
[0049] In some implementations, as shown in the figure, platform 220 may be hosted in a cloud computing environment 222. It should be noted that although the implementations described herein describe platform 220 as being hosted in cloud computing environment 222, in some implementations, platform 220 is not cloud-based (i.e., it can be implemented outside of a cloud computing environment) or may be partially cloud-based.
[0050] Cloud computing environment 222 includes an environment that hosts platform 220. Cloud computing environment 222 may provide services such as computing, software, data access, storage, etc., without the end user (e.g., user device 210) knowing the physical location and configuration of the systems and / or devices that host platform 220. As shown in the figure, cloud computing environment 222 may include a collection of computing resources 224 (collectively referred to as "computing resources 224", and a single computing resource is referred to as "computing resource 224").
[0051] Computing resources 224 include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some implementations, computing resources 224 may control platform 220. Cloud resources may include computing instances running in computing resources 224, storage devices provided in computing resources 224, data transmission devices provided by computing resources 224, etc. In some implementations, computing resources 224 may communicate with other computing resources 224 via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.
[0052] Further, as Figure 2As shown, the computing resources 224 include a collection of cloud resources, such as one or more applications ("APP") 224-1, one or more virtual machines ("VM") 224-2, virtualized memory ("VS") 224-3, one or more hypervisors ("HYP") 224-4, etc.
[0053] The application 224-1 includes one or more software applications that can be provided to and / or accessed by the user device 210 and / or the sensor device 220. The application 224-1 can eliminate the need to install and run software applications on the user device 210. For example, the application 224-1 can include software associated with the platform 220 and / or any other software that can be provided through the cloud computing environment 222. In some implementations, one application 224-1 can send information to and / or receive information from one or more other applications 224-1 through the virtual machine 224-2.
[0054] The virtual machine 224-2 includes a software implementation of a machine (e.g., a computer) that runs programs, similar to a physical machine. Depending on the degree of correspondence and use of the virtual machine 224-2 to any actual machine, the virtual machine 224-2 can be a system virtual machine or a process virtual machine. The system virtual machine can provide a complete system platform that supports the operation of a full operating system ("OS"). The process virtual machine can run a single program and can support a single process. In some implementations, the virtual machine 224-2 can run on behalf of a user (e.g., the user device 210) and can manage the infrastructure of the cloud computing environment 222, such as data management, synchronization, or long-term data transfer.
[0055] The virtualized memory 224-3 includes one or more storage systems and / or one or more devices that use virtualization technology within the storage system or device of the computing resources 224. In some implementations, in the context of the storage system, the types of virtualization can include block virtualization and file virtualization. Block virtualization can refer to the abstraction (or separation) of logical storage from physical storage, such that the storage system can be accessed without considering the physical storage or heterogeneous structure. The separation can allow the administrator of the storage system to have flexibility in how the administrator manages the storage of end users. File virtualization can eliminate the dependence between the data accessed at the file level and the location where the file is physically stored. This can optimize memory usage, server consolidation, and / or non-disruptive file migration performance.
[0056] The hypervisor 224-4 can provide hardware virtualization technology that allows multiple operating systems (e.g., "guest operating systems") to run simultaneously on a host computer such as computing resource 224. The hypervisor 224-4 can present a virtual operating platform to the guest operating systems and can manage the operation of the guest operating systems. Multiple instances of each operating system can share the virtualized hardware resources.
[0057] The network 230 includes one or more wired networks and / or wireless networks. For example, the network 230 can include a cellular network (e.g., a fifth-generation (5G) network, a long-term evolution (LTE) network, a third-generation (3G) network, a code-division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, etc., and / or a combination of these or other types of networks.
[0058] Figure 2 The number and arrangement of the devices and networks shown are provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks arranged differently from those Figure 2 shown. Additionally, Figure 2 two or more of the devices shown can be implemented within a single device, or Figure 2 a single device shown can be implemented as multiple distributed devices. Additionally or alternatively, a set of devices (e.g., one or more devices) in environment 200 can perform one or more functions described as being performed by another set of devices in environment 200.
[0059] Figure 3 is a diagram of example components of device 300. Device 300 can correspond to user device 210 and / or platform 220. As Figure 3 shown, device 300 can include a bus 310, a processor 320, a memory 330, a storage component 340, an input component 350, an output component 360, and a communication interface 370.
[0060] The bus 310 includes components that allow communication between the components of the device 300. The processor 320 is implemented in hardware, firmware, or a combination of hardware and software. The processor 320 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some implementations, the processor 320 includes one or more processors that can be programmed to perform functions. The memory 330 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by the processor 320.
[0061] The storage component 340 stores information and / or software related to the operation and use of the device 300. For example, the storage component 340 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cassette tape, a magnetic tape, and / or another type of non-transitory computer-readable medium, and a corresponding drive.
[0062] The input component 350 includes components that allow the device 300 to receive information, such as information input through user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone). Additionally or alternatively, the input component 350 may include sensors for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator). The output component 360 includes components that provide output information from the device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs)).
[0063] The communication interface 370 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter) that enable the device 300 to communicate with other devices, such as by a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. The communication interface 370 may allow the device 300 to receive information from another device and / or provide information to another device. For example, the communication interface 370 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.
[0064] Device 300 may perform one or more of the processes described herein. Device 300 may perform these processes in response to a processor 320 running software instructions stored by a non-transitory computer-readable medium such as a memory 330 and / or a storage component 340. As used herein, a computer-readable medium is defined as a non-transitory memory device. A memory device includes a memory space within a single physical storage device or a memory space distributed across multiple physical storage devices.
[0065] The software instructions may be read into the memory 330 and / or the storage component 340 from another computer-readable medium through a communication interface 370, or from another device into the memory 330 and / or the storage component 340. When run, the software instructions stored in the memory 330 and / or the storage component 340 may cause the processor 320 to perform one or more of the processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with the software instructions to perform one or more of the processes described herein. Accordingly, the implementations described herein are not limited to any specific combination of hardware circuitry and software.
[0066] Figure 3 The number and arrangement of the components shown are provided by way of example. In practice, device 300 may include additional components, fewer components, different components, or components arranged differently than those Figure 3 shown. Additionally or alternatively, a set of components of device 300 (e.g., one or more components) may perform one or more functions described as being performed by another set of components of device 300.
[0067] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Modifications and variations may be made in light of the above disclosure, or may be obtained from practice of the implementations.
[0068] As used herein, the term “component” is intended to be broadly construed as hardware, firmware, or a combination of hardware and software.
[0069] It is apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual special control hardware or software code used to implement these systems and / or methods does not limit the implementations. Accordingly, the operations and behavior of the systems and / or methods are not described herein with reference to a particular software code—it should be understood that the software and hardware may be designed to implement the systems and / or methods based on the description herein.
[0070] Even if specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may only directly depend on one claim, the disclosure of possible implementations includes the combination of each dependent claim with every other claim in the claim set.
[0071] Elements, acts, or instructions used herein should not be construed as critical or essential unless explicitly described as such. Additionally, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Further, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." Where only one item is intended, the term "a" or similar language is used. Additionally, as used herein, the terms "having," "comprising," "containing," or similar terms are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise explicitly stated.
Claims
1. A method for performing information retrieval using sentence similarity, comprising: Receiving, by a device, a first sentence including a first set of words; Receiving, by the device, a second sentence including a second set of words; Generating, by the device and using a word embedding model, a first set of vectors corresponding to the first set of words of the first sentence; Generating, by the device and using the word embedding model, a second set of vectors corresponding to the second set of words of the second sentence; Determining, by the device based on the first set of vectors and the second set of vectors, a cosine similarity of words of the first sentence and the second sentence; Calculating, by the device based on the first set of vectors, the second set of vectors, and the cosine similarity, a Needleman-Wunsch matrix; Determining, based on the Needleman-Wunsch matrix, a sentence similarity between the first sentence and the second sentence, the Needleman-Wunsch matrix storing alignment scores of alignments of the first set of vectors and the second set of vectors, the sentence similarity being determined based on the alignment scores; and Sending, by the device, the sentence similarity to allow information retrieval.
2. The method according to claim 1, wherein, The calculating the Needleman-Wunsch matrix includes: using the Needleman-Wunsch technique to generate the Needleman-Wunsch matrix.
3. The method according to claim 1, wherein, The word embedding model is a word2vec model.
4. The method according to claim 1, further comprising: Starting from the lower right cell of the Needleman-Wunsch matrix, comparing a set of alignment scores associated with a set of directions of the Needleman-Wunsch matrix to determine an alignment score provided by a channel through the Needleman-Wunsch matrix.
5. The method according to claim 4, wherein, The set of directions includes a top direction, a left direction, and a left diagonal direction based on the lower right cell of the Needleman-Wunsch matrix.
6. The method according to claim 4, further comprising: Determining a sentence similarity between the first sentence and the second sentence based on the alignment scores.
7. A device, comprising: At least one memory configured to store program code; At least one processor configured to read the program code and operate according to instructions of the program code, the program code including: A receiving code configured to cause the at least one processor to receive a first sentence including a first set of words and receive a second sentence including a second set of words; A generating code configured to cause the at least one processor: use a word embedding model to generate a first set of vectors corresponding to the first set of words of the first sentence, use the word embedding model to generate a second set of vectors corresponding to the second set of words of the second sentence; determine, by the device based on the first set of vectors and the second set of vectors, a cosine similarity of words of the first sentence and the second sentence; calculate, by the device based on the first set of vectors, the second set of vectors, and the cosine similarity, a Needleman-Wunsch matrix; Determination code, configured to cause the at least one processor to determine the sentence similarity between the first sentence and the second sentence based on the Needleman-Wunsch matrix, the Needleman-Wunsch matrix storing alignment scores for the alignment of the first vector set and the second vector set, and the sentence similarity being determined based on the alignment scores; and Transmission code, configured to cause the at least one processor to transmit the sentence similarity to enable information retrieval.
8. The apparatus according to claim 7, wherein The generation code is further configured to cause the at least one processor to use the Needleman-Wunsch technique to generate the Needleman-Wunsch matrix.
9. The device according to claim 7, wherein, The word embedding model is a word2vec model.
10. The apparatus according to claim 7, further comprising: Comparison code, configured to cause the at least one processor to start from the bottom-right cell of the Needleman-Wunsch matrix and compare a set of alignment scores associated with a set of directions of the Needleman-Wunsch matrix to determine the alignment scores provided through the channels of the Needleman-Wunsch matrix.
11. The device according to claim 10, wherein, The set of directions includes a top direction, a left direction, and a left diagonal direction based on the bottom-right cell of the Needleman-Wunsch matrix.
12. The apparatus according to claim 10, further comprising: Identification code, configured to cause the at least one processor to determine the sentence similarity between the first sentence and the second sentence based on the alignment scores.
13. A non-transitory computer-readable medium storing instructions, the instructions including one or more instructions which, when run by one or more processors of a device, cause the one or more processors to perform the method according to any one of claims 1 to 6.
14. An apparatus, comprising: A receiving unit, configured to receive a first sentence including a first set of words and receive a second sentence including a second set of words; A generating unit, configured to use a word embedding model to generate a first vector set corresponding to the first set of words of the first sentence, use the word embedding model to generate a second vector set corresponding to the second set of words of the second sentence; and determine, by the apparatus, the cosine similarity of the words of the first sentence and the second sentence based on the first vector set and the second vector set; Calculate, by the apparatus, a Needleman-Wunsch matrix based on the first vector set, the second vector set, and the cosine similarity; A determination unit, configured to determine the sentence similarity between the first sentence and the second sentence based on the Needleman-Wunsch matrix, the Needleman-Wunsch matrix storing alignment scores for the alignment of the first vector set and the second vector set, and the sentence similarity being determined based on the alignment scores; And A transmitting unit, configured to transmit the sentence similarity to enable information retrieval.
Citation Information
Patent Citations
Document alignment systems for legacy document conversions
US20070150443A1
Full text query and search systems and method of use
US20110055192A1