Computer implementation methods, computer programs, computer systems (specificity ranking of text elements and its applications)

By calculating specificity scores based on contextual distances in embedding spaces, the method accurately ranks text elements, improving performance and resource efficiency in information extraction.

JP7893570B2Active Publication Date: 2026-07-22INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2022-08-22
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Conventional techniques for estimating text element specificity rely solely on metrics derived from pre-trained embedding vectors, lacking contextual analysis, leading to inaccurate specificity scores.

Method used

A method that calculates embedding vectors for text elements, selects context fragments, and determines specificity scores based on distances within the embedding space, incorporating contextual information to rank text elements by their homogeneity of appearance.

Benefits of technology

Provides accurate specificity scores that improve performance and resource efficiency in information extraction applications by distinguishing highly specific from general terms, reducing processing resources and enhancing data extraction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893570000003
    Figure 0007893570000003
  • Figure 0007893570000004
    Figure 0007893570000004
  • Figure 0007893570000005
    Figure 0007893570000005
Patent Text Reader

Abstract

To rank a plurality of text elements, each comprising at least one word, by specificity.SOLUTION: For each text element to be ranked, a method includes computing an embedding vector that locates a text element in an embedding space, and selecting a set of text fragments from reference text. Each of these text fragments contains the text element to be ranked and further text elements. For each text fragment, respective distances in the embedding space between the further text elements are calculated. The method further includes calculating a specificity score for the text element to be ranked, and storing the specificity score. After ranking the plurality of text elements, a text data structure using the specificity scores for the text elements is processed to extract data having a desired specificity from the data structure.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Statement by the inventor or co-inventor regarding prior disclosure] The following disclosure is filed under § 102(b)(1)(A) of the United States Patent Act: Certain functions of the Disclosure, designed by Francesco Fusco and Peter Willem Jan Staar, are stored on servers of the assignee of this patent application, and are available as a service through the IBM Research Deep Search platform as of March 2021.

[0002] This invention relates, in general terms, to the specificity ranking of text elements. Computer implementations for ranking multiple text elements by specificity are provided, along with applications of such methods. Systems and computer program products for implementing these methods are also provided. [Background technology]

[0003] The specificity of text elements, such as words or phrases, is a measure of the amount of information contained within those elements. If a text element contains a lot of information within a given domain, it is highly specific to that domain, and vice versa. Text specificity has been estimated in the context of search systems to evaluate whether a search query should return general or specific results, or to suggest alternative search queries to the user. Most conventional techniques for estimating specificity use statistics based on an analysis of part of the speech (e.g., how often nouns are modified) or the frequency of occurrence of specific terms. One technique evaluates the specificity of terms using various metrics derived from vectors that locate terms in an embedding space generated via a word embedding scheme. This technique utilizes metrics obtained by analyzing the distribution of embedding vectors in pre-trained embeddings. Once the embedding matrix is ​​trained, the vector distribution in the embedding space is the sole factor used to evaluate specificity. [Overview of the project] [Problems that the invention aims to solve]

[0004] Most conventional techniques for estimating specificity use statistics based on an analysis of a portion of the speech (e.g., how often nouns are modified) or the frequency of occurrence of specific terms. One technique evaluates the specificity of terms using various metrics derived from vectors that locate terms in an embedding space generated via a word embedding scheme. This technique leverages metrics obtained by analyzing the distribution of embedding vectors in pre-trained embeddings. Once the embedding matrix is ​​trained, the vector distribution in the embedding space is the only factor used to evaluate specificity. [Means for solving the problem]

[0005] One aspect of the present invention provides a computer-implemented method for ranking a plurality of text elements each containing at least one word by specificity. For each text element to be ranked, the method includes calculating an embedding vector that locates the text element in an embedding space via a word embedding scheme, and selecting a set of text fragments from a reference text. Each of these text fragments includes the text element to be ranked and a further text element. For each text fragment, the method calculates the respective distance in the embedding space between the further text element, which is each located in the space by the embedding vector calculated via the word embedding scheme, and the text element to be ranked. The method further includes calculating a specificity score for the text element to be ranked depending on the above-mentioned distance, and storing the specificity score. The resulting specificity scores for the plurality of text elements define the ranking of the text elements by specificity.

[0006] Each further embodiment of the present invention provides a computing system adapted to implement a method for ranking text elements as described above, and a computer program product including a computer-readable storage medium embodying program instructions executable by the computing system to cause the computing system to implement such a method.

Brief Description of the Drawings

[0007] Embodiments of the present invention will be described in more detail below by way of illustrative and non-limiting examples with reference to the accompanying drawings.

[0008] [Figure 1] It is a schematic representation of a computing system implementing a method for embodying the present invention. [Figure 2] It is a diagram showing component modules of a system embodying the present invention for ranking text elements by specificity.

[0009] [Figure 3] A diagram showing steps of a ranking method executed by the system shown in FIG. 2.

[0010] [Figure 4] A diagram showing component modules of a text element ranking system in an embodiment of the present invention.

[0011] [Figure 5] A diagram showing steps of a word embedding process in the system shown in FIG. 4.

[0012] [Figure 6] A schematic diagram of a word embedding process.

[0013] [Figure 7] A diagram showing the operation of a context fragment selector in the system shown in FIG. 4.

[0014] [Figure 8] A diagram showing steps of a specificity score calculation process in the system shown in FIG. 4.

[0015] [Figure 9A] A diagram showing a specificity ranking obtained in one implementation of the system shown in FIG. 4. [Figure 9B] A diagram showing a specificity ranking obtained in one implementation of the system shown in FIG. 4.

[0016] [Figure 10] A diagram showing the operation steps of an application using the text element ranking method embodying the present invention. [Figure 11] A diagram showing the operation steps of an application using the text element ranking method embodying the present invention. [Figure 12]This diagram shows the operational steps of an application using the text element ranking method that embodies the present invention. [Figure 13] This diagram shows the operational steps of an application using the text element ranking method that embodies the present invention. [Modes for carrying out the invention]

[0017] Some embodiments of the present invention may be systems, methods, or computer program products or combinations thereof. A computer program product may include a computer-readable storage medium (or a set of mediums) having computer-readable program instructions that cause a processor to execute aspects of the present invention.

[0018] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of those described above. A non-exhaustive list of more specific examples of computer-readable storage media includes, namely, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital multipurpose disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or grooved raised structures recording instructions, and any suitable combination of those described above. When used herein, computer-readable storage media should not be interpreted as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or transient signals such as electrical signals transmitted through wires.

[0019] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives computer-readable program instructions from the network and transfers such computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.

[0020] The computer-readable program instructions that perform the operation of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, where one or more programming languages ​​include object-oriented programming languages ​​such as Smalltalk®, C++, etc., and conventional procedural programming languages ​​such as the C programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or fully on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), and the connection may be to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may be personalized by executing computer-readable program instructions using state information of computer-readable program instructions in order to perform an aspect of the present invention.

[0021] Aspects of the present invention are described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block in a flowchart or block diagram, or both, and any combination of blocks in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.

[0022] These computer-readable program instructions may be provided to the processor of a general-purpose computer, a dedicated computer, or another programmable data processing device to generate a machine, thereby creating means for implementing functions / operations specified in one or more blocks of a flowchart or block diagram, or both, through the instructions executed via the processor of the computer or other programmable data processing device. Furthermore, these computer-readable program instructions may be stored in a computer-readable storage medium, which can instruct a computer, a programmable data processing device, or another device, or a combination thereof, to function in a specific manner, thereby including a product containing instructions that implement modes of functions / operations specified in one or more blocks of a flowchart or block diagram, or both.

[0023] Furthermore, computer-readable program instructions may be loaded into a computer, another programmable data processing device, or another device to execute a series of operational steps on the computer, another programmable device, or another device, thereby generating a computer implementation process in which the instructions executed on the computer, another programmable device, or another device implement the functions / operations specified in one or more blocks of a flowchart or block diagram, or both.

[0024] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions containing one or more executable instructions that implement a specified logical function. In some alternative implementations, the functions described in a block may be performed in an order different from the order shown in the drawings. For example, two blocks shown consecutively may actually be executed substantially simultaneously, and blocks may be executed in reverse order depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, or both, and any combination of blocks in a block diagram or flowchart, or both, may be implemented by a dedicated hardware-based system that performs a specified function or operation, or a combination of dedicated hardware and computer instructions.

[0025] The embodiments to be described can be implemented as a computer implementation method for ranking text elements by specificity. Such a method may be implemented by a computing system comprising one or more general-purpose or dedicated computers that provide the functionality to implement the operations described herein, each of which may comprise one or more (actual or virtual) machines. Steps of the method embodying the present invention may be implemented by program instructions, such as program modules, implemented by the processing unit of the system. Generally, a program module may include routines, programs, objects, components, logic, data structures, etc., that perform a particular task or implement a particular abstract data type. The computing system may be implemented in a distributed computing environment, such as a cloud computing environment, in which tasks are performed by remote processing devices linked through a communication network. In a distributed computing environment, program modules may reside in both local computer system storage media, including memory storage devices, and remote computer system storage media.

[0026] Figure 1 is a block diagram of an exemplary computing device that implements a method embodying the present invention. The computing device is shown in the form of a general-purpose computer 1. The components of computer 1 may include processing units such as one or more processors represented by processing units 2, system memory 3, and a bus 4 that connects various system components, including the system memory 3, to the processing units 2.

[0027] Bus 4 represents one or more of several types of bus structures, including memory buses or memory controllers, peripheral buses, accelerated graphics ports, and processor or local buses using any of the various bus architectures. Such architectures, but are not limited to, include the Industrial Standard Architecture (ISA) bus, Microchannel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0028] Computer 1 typically includes a variety of computer-readable media. Such media may be any available media accessible by Computer 1, including volatile and non-volatile media, and removable and non-removable media. For example, system memory 3 may include computer-readable media in the form of volatile memory such as random access memory (RAM) 5 or cache memory 6 or both. Computer 1 may further include other removable / non-removable, volatile / non-volatile computer system storage media. For illustrative purposes only, a storage system 7 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium (commonly called a “hard drive”). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a “floppy disk”), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM, or other optical media may also be provided. In such cases, each may be connected to bus 4 by one or more data media interfaces.

[0029] Memory 3 may include at least one program product having one or more program modules configured to perform the functions of embodiments of the present invention. For example, a program / utility 8 having a set (at least one) of program modules 9 may be stored in Memory 3, as may an operating system, one or more application programs, other program modules, and program data. Each of these, or any combination thereof, may include an implementation of a networking environment. Program modules 9 generally perform the functions or methodologies of embodiments of the present invention, or both, as described herein.

[0030] Furthermore, computer 1 may communicate with one or more external devices 10 such as a keyboard, pointing device, display 11, one or more devices that enable a user to interact with computer 1, or any device that enables computer 1 to communicate with one or more other computing devices (e.g., a network card, modem, etc.), or a combination thereof. Such communication may occur via an input / output (I / O) interface 12. Computer 1 may also communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter 13. As shown in the figure, the network adapter 13 communicates with other components of computer 1 via bus 4. Computer 1 may also communicate with additional processing units 14, such as a GPU (graphics processing unit) or FPGA, that implement embodiments of the present invention. It should be understood that other hardware and / or software components, or both, may be used in conjunction with computer 1, although these are not shown. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.

[0031] Figure 2 schematically shows component modules of an exemplary computing system embodying the present invention. System 20 comprises memory 21 and control logic, collectively shown as 22, which includes the function of ranking text elements by specificity. The control logic 22 includes a word embedding module 23, a context selector module 24, and a specificity calculator module 25. Each of these modules includes the function of implementing a particular stage of the ranking process, which is detailed below. These modules interface with memory 21, which stores various data structures used in the operation of system 20. These data structures include a set of N text elements 27 (here {t i This is shown by {t} and includes a set of embedding vectors 28 generated by the word embedding module 23 and a set of text fragments ("context fragments") 29 selected by the context selector 24 from the reference text represented by the text corpus 30 in the drawing. i A set of singularity scores 31 (here {S i This is shown by}, and i=1~N) is also stored in system memory 21.

[0032] Generally, the functions of logic modules 23-25 ​​may be implemented by software (such as program modules), hardware, or a combination thereof. The functions described may be assigned differently among system modules in other embodiments, and the functions of one or more modules may be combined. The component modules of system 20 may be provided in one or more computers of the computing system. For example, all modules may be provided in computer 1, or the modules may be provided in one or more computers / servers to which a user computer can connect via a network (which may include one or more component networks or an internetwork (including the Internet) or both) for inputting text items to be ranked. System memory 21 may be implemented by one or more memory / storage components associated with one or more computers of system 20.

[0033] Set of text elements {t i {t} may include individual words or multiword expressions (MWE) or both, and may be compiled for a specific application / domain or may span multiple domains for use in various applications. Some embodiments of the present invention utilize the intrinsic specificity of these elements to take advantage of {t i Incorporate MWE into the}. MWE t i The list can be pre-compiled either manually or automatically, as described below, for storage in system memory 21.

[0034] The reference text corpus 30 may be local or remote from system 20, and the elements to be ranked will be {t iIt may include text from one or more information sources spanning the region of {}. Although represented as a single entity in FIG. 2, the reference text corpus may include content distributed across multiple information sources, such as databases or websites or both, which the system may dynamically access via a network. In some embodiments, the reference text 28 may be pre-compiled for system operation and stored in the system memory 21.

[0035] The flowchart 300 of FIG. 3 shows the steps of a ranking process executed by the system 20 (designated as the text element ranking step 34). Step 35 represents the storage of the set of text elements {t to be ranked in the system memory 21. In step 36, the word embedding module 23 embeds the text element t via a word embedding scheme i} into the system memory 21. iAn embedding vector is calculated for each element. Word embedding schemes are well known and essentially generate a mapping between text elements and real-valued vectors that define the location of each text element in a multidimensional embedding space. The relative locations of text elements in this space indicate the degree of relationship between them, with "closer" elements in the embedding space being more closely related than further-placed elements. In particular, the concept of word embedding is to map elements that appear in similar text contexts to be close to each other in the embedding space. Any desired word embedding scheme may be used here, including context-independent schemes using Word2Vec, GloVe (Global Vectors), or FastText models, or context-dependent schemes including transformer architecture-based models such as BERT (Bidirectional Encoder Representations from Transformers) models. A word embedding scheme can generate a "cloud" of text elements that appear in similar contexts and therefore represent semantically similar concepts. Each embedding vector generated by the word embedding module 23 thus corresponds to the text element t in the embedding space, indicated here by χ. i The location is determined. The resulting vector is stored in 28 in system memory 21.

[0036] In stage 37, the context selector module 24 determines the element t that will be ranked. i Select a set of text fragments from reference text 30 for each. Each of these text fragments is element t i And further text elements (where these further text elements will be ranked by other elements t i (may include or not include one or more of the above). For example, the context selector 24 includes element t in the reference text. iA few lines of text, a sentence or paragraph containing, or a given element t within the text i You may select a window of words surrounding the given text element. Generally, one or more text fragments containing a given text element may be selected here, and in some embodiments, multiple fragments are selected for each element. The selected text fragments are stored in system memory 21 as context fragments 29 (possibly after further processing as described below).

[0037] Stages 38-40 demonstrate the operation of the singularity calculator 25. These stages represent the ranking of the text elements t. i This is performed for each step. In step 38, the singularity calculator takes a given element t from the set of context fragments 29. i Extract the context fragment containing . For each fragment, the singularity calculator 25 calculates the element t in the embedding space χ. i Then, the distance between each of the further text elements within that fragment is calculated. To calculate these distances, each further text element must first be located in the embedding space χ by a corresponding embedding vector calculated via a word embedding scheme. The embedding vectors for the further text elements may be pre-calculated in step 36 above, for example, via a context-independent embedding scheme as detailed below, or they may be calculated dynamically via a context-dependent embedding scheme. t in χ i The distance between element t and further text elements can be conveniently calculated as the cosine similarity between the two vectors representing these elements. However, in other embodiments, any convenient distance metric such as the Euclidean distance may be used. In step 39, the singularity calculator calculates the distance between elements t i Calculate the specificity score for element t. i Specificity score S for iThis depends on the distance calculated in step 38 from the context fragment containing the element. The singularity score may be calculated from these distances in various ways, as described below. In step 40, the singularity score S i It is stored in set 31 in system memory 21. All text elements {t i After processing the context fragments for}, the set obtained as a result of the specificity score is {S i Next, the ranking of these text elements is determined by their specificity.

[0038] The method described above corresponds to the context of text elements in calculating specificity by using the distance between text elements within context fragments in the embedding space χ. By injecting context into the information extracted from word embeddings, the resulting specificity score provides a measure of the homogeneity of the context in which the text elements appear. This provides a true measure of specificity, based on the fact that highly specific terms tend to appear in more homogeneous contexts than more general terms. As an exemplary example, the term "information hiding" is a highly technical expression used in software engineering when designing data structures that do not expose internal state to the outside. In contrast, "hiding information" is a term that can and will appear in many different contexts. Therefore, the technique described above provides an improved specificity estimation with resulting advantages in terms of performance and resource efficiency of information extraction applications. This technique is also entirely unsupervised, making it possible to calculate specificity scores for any set of text elements without requiring annotated training data.

[0039] Diagram 400 in Figure 4 shows a more detailed system implementation in several embodiments of the present invention. System 45 in this embodiment is {m iThe system is adapted to compile and rank a large set of MWEs, as shown by}. The control logic 46 of this system includes a word embedding module 47, a context selector 48, and a specificity calculator 49, as previously described. The control logic also includes an MWE extractor module 50 and a text encoder module 51. The data structure stored in the system memory 53 is a set of MWEs 54{m} that are automatically compiled by the MWE extractor 50 from a knowledge base schematically shown in 55. i This includes a knowledge base 55 and a tokenized text dataset 56 generated by a text encoder 51 from a text corpus designated as a WE (word embedding) corpus 57. In practice, the knowledge base 55 and the WE corpus 57 may represent content collected from multiple sources or distributed across multiple sources. Memory 53 also stores an embedding matrix 58 generated by a word embedding module 47 and a set of inverse frequencies 59, which will be further described below. In addition, memory 53 stores a set of context fragments 60 generated by a context selector 48 and a set of instance scores 61, which will be further described below, in MWE{m i Store this along with the final set of 62 specificity scores calculated for}.

[0040] The operation of system 45 will be explained with reference to Figures 5 to 8. The flowchart 500 in Figure 5 shows the steps leading to the generation of the embedding matrix 58. In step 65, the MWE extractor 50 accesses the knowledge base 55 and extracts MWE associated with hyperlinks within the knowledge base. A knowledge base (such as Wikipedia, DBPedia, Yago, etc.) is essentially a graph of concepts where concepts are linked to one another. The MWE extractor 50 can extract MWE from the knowledge base by searching through hyperlinks. For example, in the following sentence (hyperlinks are indicated by underlines): "In thermal power stations, mechanical power is produced by a heat engine which converts thermal energy, from combustion of a fuel, into rotational energy," the MWE extractor may select "heat engine" and "thermal energy." Hyperlinks in such knowledge bases are manually annotated and therefore of high quality. By simply scanning the knowledge base text, the MWE extractor 50 can extract a vast number of well-defined MWEs. In this example, the MWE extractor explores the knowledge base 55 to compile a large dictionary of MWEs covering a wide range of topics. The resulting set of MWEs {m i}54 is stored in memory 53 in step 66.

[0041] In steps 67 and 68, the text encoder 51 processes each MWE m within the tokenized text. iTokenized text 56 is generated by preprocessing and tokenizing the WE corpus 57 so that each word is encoded as a single token and other words in the corpus are encoded as their respective tokens. In particular, in step 67, the text encoder preprocesses the WE corpus 57 as schematically shown in the data flow diagram 600 of Figure 6. i Instances of are identified within the raw corpus, and each of these is concatenated and treated as an individual word. For example, the MWE "machine learning" is concatenated as "machine_learning". All units and stop words (such as "a", "and", "was", etc.) are also removed during preprocessing, and all uppercase letters are changed to lowercase. The resulting text is then split into sentences to train word embeddings. In step 68 of Figure 5, the preprocessed text is tokenized by encoding all remaining words and MWEs as their respective single tokens. One-hot encoding is used here for convenience, but other encoding schemes can certainly be assumed. Thus, each token represents a specific word / MWE, and whenever that word / MWE appears in the preprocessed text, it is replaced with the corresponding token. The resulting tokenized text 56 is stored in system memory 53.

[0042] In step 69, the word embedding module 47 processes the tokenized text 56 to generate an embedding matrix 58. In this embodiment, the tokenized sentences in text 56 are used to train the Word2Vec embedding model using known CBOW (Common Bag Of Words) and negative sampling techniques (see, e.g., "Distributed representations of words and phrases and their compositionality," Mikolov et al., Advances in Neural Information Processing Systems 26, 2013, pp. 3111-3119). As a result, a set of embedding vectors is obtained, one for each token corresponding to each MWE / word in the preprocessed text, as schematically shown in Figure 6. This set of vectors constitutes the embedding matrix 58, which is stored in system memory 53 in step 70. Thus, the resulting embedding matrix contains embedding vectors corresponding to the text elements to be ranked (here, MWEs) and further text elements contained within the text fragments selected by the context selector 48 from the reference text corpus 30. (In this regard, a separate reference text corpus 30 is shown in Figure 4, but in other embodiments, the WE corpus 57 may serve as the reference text for the context fragment, thereby making the embedding vector available for all text elements within the context fragment.)

[0043] When processing the WE corpus 57, the embedding module 47 counts the number of instances of each text element (in this case, MWE or word, represented collectively by w) in the preprocessed corpus. For each element w, the embedding module calculates its inverse frequency f(w). The inverse frequency of an element w that appears n times in a corpus of m words is defined as f(w) = m / n. The set of inverse frequencies f(w) for each element w is stored in 59 in system memory 53.

[0044] The operation of the context selector 48 is shown here in the data flow of diagram 700 in Figure 7. In this example, the context selector 48 uses a reference text corpus 30 which is separate from the WE corpus 57. In stage (a) of Figure 7, the context selector extracts sentences from the reference corpus 30. In stage (b), the MWE m in the sentence i All instances of are identified and marked as shown by bold and linked in the drawing. All common stop words, units and numbers are also identified as shown by strikethrough in stage (b), and these are removed in stage (c) to obtain the processed sentence. MWE m i Each processed statement containing an instance of is selected as a context fragment. The context selector then stores each context fragment in set 60, here as a "bag of word" (BOW), as shown in step (d).

[0045] The flowchart 800 in Figure 8 illustrates the operation of the singularity calculator 49 in this embodiment. In step 75, the singularity calculator selects a context fragment from the fragment set 60. Subsequent steps 76-78 then calculate the MWE in the BOW for the selected fragment. i This is performed each time. In step 76, the singularity calculator calculates the MWE m in the BOW in the embedding space χ, where the embedding vector is contained within the embedding matrix 58. iThe distance between this and each further text element (MWE / word) w is calculated. Here, d(m i The distances shown by ,w) are, respectively, m i It is calculated as the cosine similarity between two vectors representing w. This yields a number within the range (-1, +1), where a higher number indicates a closer element m in the embedding space χ. i And w is shown.

[0046] The singularity calculator then calculates the MWE m from the distance calculated in step 76 for the current fragment. i The instance score 61 is calculated for each distance d(m). i ,w) is first weighted based on the inverse frequency f(w) stored in set 59 for element w, and MWE m i The instance score for is calculated as a function of the weighted distance for each fragment. In particular, in step 78, the singularity calculator obtains the instance score by aggregating the weighted distances for each fragment. In this example, MWE m i and further elements w1, ..., w k Given a BOW containing the following, the instance score T i but,

number

[0047] If further context fragments are to be processed in decision stage 79, the operation returns to stage 75, where the next fragment is selected from set 60 and processed as described above. If all context fragments have been processed in stage 79, the operation proceeds to stage 80. Here, MWE m i For each, the singularity calculator 49 calculates m i Instance score T i The specificity score S as a function of i The specificity score S is calculated. In this embodiment, the specificity score S i Here, the simple average is:

number

[0048] By weighting the distance by the inverse frequency described above, the contribution of common (and possibly more common) elements w is penalized, and a bias is introduced that directs the mean towards less common (and possibly more singular) elements. By generating an embedding matrix 58 from a large and diverse WE corpus 57 with a large dictionary of MWEs, the above system can automatically generate specificity scores for use in a wide range of applications. However, in general, the specificity score {S i The} may be computed for any subset of tokens in the embedding space χ for MWEs or individual words or both, and this subset may be specific to a given field or application. In other embodiments, the embedding matrix 58 may also be generated for MWEs / words related to specific technical fields / applications.

[0049] The tables in Figure 9A, i.e., 900a, and Figure 9B, i.e., 900b, show extractions from specificity rankings generated by the implementation of the system in Figure 4. The results in Figure 9A were obtained using a reference text corpus containing 1.5 million patent abstracts. The results in Figure 9B were obtained using a reference text corpus containing 1.2 million abstracts from arXiv articles. Both sets of results used embedding matrices constructed from a WE corpus of over 100 million news articles. Figure 9A shows specificity scores for 10 highest-scoring MWEs and 10 lowest-scoring MWEs with tokens containing the word "knowledge". Figure 9B shows 10 highest-scoring MWEs and 10 lowest-scoring MWEs with tokens containing the word "language". The scores may correlate well with the specificity of the enumerated MWEs. These examples demonstrate that specificity scores calculated by the techniques described above can reliably distinguish highly technical MWEs from more general representations, even as a simple average of instance scores calculated across a large reference corpus.

[0050] The specificity ranking technique can be used to improve the performance of numerous data processing applications. In this technique, after ranking text elements by specificity, the specificity score is used when processing the text data structure to extract data with a desired specificity. The use of specificity scores can reduce the processing resources required to extract relevant data from various data structures for various purposes, improve the quality of the extracted data, and thus improve the performance of applications using these data structures. Several exemplary applications are described below with reference to Figures 10 to 13.

[0051] Flowchart 1000 in Figure 10 illustrates the operation of a knowledge guidance system for extracting knowledge from a large text corpus. Such applications typically process vast amounts of text mined from databases / websites in the cloud. Step 85 represents the storage of the cloud data to be analyzed. In step 86, the ranking method described above is used to rank the text elements in this data by specificity. In step 87, the cloud data is filtered to identify the most specific text elements in the corpus based on specificity scores, e.g., a set of text elements with specificity scores higher than a defined threshold. In step 88, a knowledge graph (KG) is then constructed from the filtered data. A knowledge graph is a well-known data structure commonly used to extract meaningful knowledge from large amounts of data for industrial, commercial, or scientific applications. Essentially, a knowledge graph contains nodes representing entities, which are interconnected by edges representing the relationships between connected entities. The knowledge graph constructed in step 88 therefore contains nodes corresponding to elements in an identified set of the most unique text elements, where the nodes are interconnected by edges representing the relationships between them. (Such relationships can be defined in various ways for a particular application, as will be apparent to those skilled in the art). The resulting knowledge graph provides a data structure that can be searched to extract the information represented in the graph. In response to an input search query in step 89, the system then searches the graph in step 90 to extract the requested data from the data structure. Filtering the data used to construct the knowledge graph in this application can significantly reduce the size of the data structure and therefore the memory required to store the graph, while at the same time ensuring that the most unique data containing the majority of the information is retained. The compute intensity of the search operation is also reduced, and the search results are narrowed down to more unique, typically more useful, information.

[0052] Another application of the specificity score relates to the extension of keyword sets for the search process. Flowchart 1100 in Figure 11 illustrates the operation of such a system. Step 95 represents the storage of a word embedding matrix into the system, which contains vectors that locate each text element in the latent embedding space. Such a matrix can be generated in a similar manner to the embedding matrix 58 in Figure 4 and may encode a wide range of words / MWEs in one or more technical fields. In step 96, the text elements in the embedding matrix are ranked by specificity as described above. Step 97 represents the input of a keyword by the user, represented by vectors in the embedding matrix relating to the field to be searched. In step 98, the system then searches the embedding space around that keyword to identify neighboring text elements in the embedding space. Various clustering / nearest neighbor search processes can be utilized here, and the search process is adapted to locate the set of most specific text elements in the neighborhood of the input keyword (e.g., elements with a specificity score above a desired threshold). In stage 99, the identified text elements are stored as an expanded keyword set, along with the user-input keywords. This expanded keyword set can then be used to explore the text corpus, for example, by string matching the keywords in the set against documents in the corpus, to identify relevant documents in the required domain. The use of specificity scores in this application allows for the automatic expansion of small user-input keyword sets with highly specific and relevant keywords, facilitating the locating of relevant documents in a given domain. A specific example of this application is collecting training documents to train a text classifier model.

[0053] Flowchart 1200 in Figure 12 illustrates the use of specificity scores in an automated phrase extraction system. Phrase extraction systems are well known and can be used to extract thematic or key phrases from documents for extraction / summarization purposes (see, for example, "Key2Vec: Automatic ranked keyphrase extraction from scientific articles using phrase embeddings," Mahata et al., Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2, June 2018, pp. 634-639). These systems often use a graph-based representation for candidate key phrases within a document. Nodes representing candidate phrases are interconnected by edges representing relationships between nodes, each with associated weights (depending on semantic similarity, frequency, etc.), which are then used to select the best candidate. Step 100 in Figure 12 represents a typical text processing operation that generates a graph for candidate phrases. In step 101, text elements in the graph are ranked by specificity using the method described above. In step 102, the graph is pruned based on the specificity scores of the text elements within the candidate phrases to obtain a subgraph representing the most specific subset of these phrases. This most specific subset may include phrases containing text elements with specificity scores exceeding a desired threshold. In step 103, the resulting subgraph is then processed in a standard manner to extract the best candidate phrases from it. Such processing may involve scoring nodes based on various graph features to extract the best phrases for the desired purpose.

[0054] Flowchart 1300 in Figure 13 illustrates the use of a specificity score in a search system. In step 105, where text elements in the search database are ranked by specificity as described above. In response to the input of a search query in step 106, the system identifies any ranked text elements in the query text. In step 108, the system generates a response to the search query by extracting data from the search database, relying on the specificity score for any ranked text elements thus identified. The response may here be to suggest an alternative search query for the user or to retrieve the requested data from the search database. The specificity score can here be used to identify the most relevant alternative query or response data based on the element with the highest specificity score in the input query. The specificity score may also be used to assess the user's level of knowledge and return results accordingly. For example, an input query containing highly specific text elements suggests that a knowledgeable user desires more detailed results, while a low-specificity query suggests that the user needs more general, high-level results.

[0055] Specificity ranking techniques can be seen to provide more efficient processing and improved results in various processing applications, and can reduce the memory and processing resources required for knowledge extraction operations.

[0056] The methods embodying this invention presuppose the understanding that highly unique text elements, such as text elements representing technical concepts, tend to appear in text contexts that are inherently homogeneous. These methods use fragments of reference text to provide context for the text elements to be ranked. In this case, the uniqueness score for a given text element is based on the distance in the embedding space between that text element and other text elements within a selected text fragment containing that element. Based on the above understanding, the methods embodying this invention construct a correspondence of context for text elements in information extracted from word embeddings such that the resulting uniqueness score provides a measure of the homogeneity of the context in which the text element appears. This provides a clearly simple technique for capturing the uniqueness of text elements, resulting in improved estimation of uniqueness and improved performance of processing systems using such estimates.

[0057] After ranking multiple text elements, a method embodying the present invention may process the text data structure using a specificity score for the text elements to extract data with a desired specificity from the data structure. Using a specificity score can reduce the processing resources required to extract relevant data from the data structure in various applications, improve the quality of the extracted data, and therefore improve performance. For example, a specificity score can be used as a filtering mechanism to reduce the memory required to store a search structure such as a knowledge graph by pruning the graph to remove unnecessary elements, thereby reducing the computational intensity of search operations performed on such graphs. Examples of other text data structures and processing applications that utilize these structures are described in more detail below.

[0058] Generally, the text elements to be ranked may include single-word text elements (i.e., individual words), multi-word expressions (i.e., text elements containing at least two words), or both. Multi-word expressions include combinations of words such as separable compound words or phrases that convey a specific meaning together or act as a semantic unit at a certain level of linguistic analysis. In some embodiments, the multiple text elements to be ranked may include multi-word expressions, taking advantage of the fact that these are often more intrinsically unique than single words. In such cases, a single embedding vector is calculated for each multi-word expression, i.e., the multi-word expression is treated as if it were a single word for the embedding process. The text elements to be ranked may, of course, include individual words and, if desired, multi-word expressions.

[0059] Some embodiments of the present invention select multiple text fragments from a reference text, each containing a text element to be ranked. For each text fragment (e.g., a sentence) containing an element to be ranked, these embodiments calculate an instance score that depends on the distance between the text element in that fragment and any further text elements. The specificity score for a text element is then calculated as a function of the instance scores for the multiple text fragments containing that element. The accuracy of the specificity score generally improves with increasing numbers of text fragments selected as context for the text element. In some embodiments of the present invention, the reference text includes a text corpus, and for each text element to be ranked, a fragment of the text corpus is selected for each instance of that text element in the corpus.

[0060] When calculating instance scores from text fragments, some embodiments weight the distance between the text element to be ranked and each further text element by the inverse frequency (described below) of that further text element in the text corpus, e.g., the corpus used to compute the embedding vectors. The instance score is calculated as a function of these weighted distances for the fragment. This weighting works to penalize contributions of more common words, giving more weight to less frequent words, and thus improving the accuracy of the specificity score.

[0061] The embedding vectors may be computed by any convenient word embedding scheme, which may include context-independent or context-dependent embedding schemes. A context-independent word embedding scheme processes a text corpus to generate an embedding matrix containing embedding vectors for selected text elements (here, words, multi-word expressions, or both) within the text. A context-dependent scheme utilizes an embedding model that can take arbitrary input text and output embedding vectors for that text. Embodiments using context-independent embeddings have been found to provide improved accuracy in specificity calculations, particularly for more technical terms. Therefore, certain methods process a text corpus to generate an embedding matrix. In particular, several embodiments tokenize the text corpus such that, within the tokenized text, each text element to be ranked is encoded as a single token, and other words in the corpus are encoded as their respective tokens. The tokenized text is then processed via a word embedding scheme to generate an embedding matrix containing embedding vectors corresponding to the text elements to be ranked and any further text elements to be extracted from text fragments selected for contextual purposes. A set of multi-word expressions, which will be encoded as multiple single tokens, can be stored before the corpus is tokenized. In some embodiments of the present invention, a set of multi-word expressions can be automatically compiled by processing a text dataset, for example, by automatically extracting expressions from a large set of documents, or by identifying hyperlinks containing multi-word expressions in text from an online knowledge base. In this way, a large dictionary of multi-word expressions can be compiled for the embedding process. All or a subset of these can then be ranked by specificity as needed.

[0062] Naturally, it will be understood that numerous changes and modifications can be made to the exemplary embodiments described. For example, the dictionary of MWEs to be ranked may be extracted by an automated phrase extraction system in other embodiments. Instance scores can be calculated in various ways by averaging, summing, or otherwise aggregating distances or weighted distances, and specificity scores may be calculated as a function of the instance scores or other underlying distances. For example, the specificity score may be based on a statistical processing of the distribution of instance scores for an element, for example, as a statistical mean after removing the highest and lowest instance scores from the distribution.

[0063] The steps in the flowchart may be implemented in a different order than shown, and some steps may be executed in parallel where appropriate. Generally, where a feature is described herein with reference to a method of embodying the invention, the corresponding feature may be provided in a computing system / computer program product that embodies the invention, and vice versa.

[0064] The descriptions of various embodiments of the present invention are presented for illustrative purposes only and are not intended to be exhaustive or limitful to the embodiments disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the embodiments described. The terminology used herein has been selected to best describe the principles of the embodiments, the practical applications of the technology found in the market or technical improvements thereto, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

Claims

1. A computer implementation method for ranking multiple text elements, The steps include: calculating an embedding vector that locates a first text element among multiple text elements that will be ranked in the embedding space via a word embedding scheme; A step of selecting a set of text fragments from a reference text, wherein each text fragment includes the first text element and at least one other text element which will be ranked; For each text fragment, the steps include: calculating the respective distance between the first text element to be ranked and at least one other text element to be ranked, which is located in the embedding space by an embedding vector calculated via the word embedding scheme in the embedding space; A step of calculating a specificity score for each text element to be ranked, depending on the respective distance in the embedding space, A step of storing the aforementioned singularity scores, wherein the singularity scores for the plurality of text elements define the ranking of the text elements by singularity, and A method that includes [a certain feature].

2. The method according to claim 1, wherein the plurality of text elements to be ranked include multi-word expressions.

3. The method according to claim 1, wherein the plurality of text elements to be ranked include single-word text elements.

4. Text corpus, Tokenizing the text corpus such that each of the text elements to be ranked is encoded as a single token, and other words in the text corpus are encoded as their respective tokens, Processing the tokenized text via the word embedding scheme to generate an embedding matrix that includes the embedding vectors corresponding to the text elements to be ranked and the at least one other text elements to be ranked. The stage processed by The method according to claim 1, further comprising:

5. The steps include: storing a set of multi-word expressions before tokenizing the aforementioned text corpus; During tokenization of the text corpus, the process involves encoding each multiword expression within the set of multiword expressions as a single token. The method according to claim 4, further comprising:

6. The method according to claim 5, further comprising the step of compiling the set of multi-word expressions by processing a text dataset.

7. For each text element that will be ranked, The steps include selecting a plurality of text fragments containing the first text element from the reference text, For each text fragment, the steps include: calculating an instance score depending on the distance between the first text element to be ranked within the text fragment and the at least one other text element; A step of calculating the specificity score as a function of the instance scores for the plurality of text fragments. The method according to claim 1, further comprising:

8. For each text fragment, A step of weighting the distance between each of the first text element and the at least one other text element, which will be ranked by the reciprocal of the frequency of occurrence of the further text elements in the text corpus, A step of calculating the instance score as a function of the weighted distance for the text fragment. The method according to claim 7, further comprising:

9. For each text element that will be ranked, The steps include: calculating an instance score for each of the multiple text fragments by aggregating the weighted distances for the aforementioned text fragments; The step of calculating the specificity score by aggregating the instance scores for the plurality of text fragments. The method according to claim 1, further comprising:

10. The aforementioned reference text includes a text corpus. The method according to claim 7, wherein for each text element to be ranked, a fragment of the text corpus is selected for each instance of the first text element in the text corpus.

11. The method according to claim 1, wherein each text fragment includes a sentence.

12. The method according to any one of claims 1 to 11, further comprising the step of ranking the plurality of text elements, processing the text data structure using the specificity scores for the text elements, and extracting data having a desired specificity from the text data structure.

13. The text data structure includes a text corpus, and the method is The steps include: constructing a knowledge graph containing the set of most unique text elements in the corpus using the uniqueness scores for the text elements in the corpus; The steps include: The method according to claim 12, further comprising:

14. The text data structure includes a word embedding matrix containing vectors that locate each text element in latent space, and the method is In response to input of text elements corresponding to a vector in the latent space, the step of identifying the set of most unique text elements in the neighborhood of the input text element in the latent space based on the uniqueness score. The method according to claim 12, further comprising:

15. The text data structure includes a graph having nodes representing text phrases, the graph being interconnected by edges representing relationships between nodes, and the method is The steps include: pruning the graph based on the specificity scores of the text elements within the text phrase to obtain a subgraph representing the most specific subset of the text phrase; The steps include processing the subgraph to extract the desired phrase and The method according to claim 12, further comprising:

16. The text data structure includes a search database, and the method is In response to the input of the search query, The steps include identifying any ranked text element within the search query, A step of generating a response to the search query by extracting data from the search database based on the specificity score for any ranked text element identified in this way. The method according to claim 12, comprising:

17. The method described above is: The steps include filtering the text elements to identify the most unique text elements that have a specificity score higher than a defined threshold, A step in which the most unique text elements are used to construct a knowledge graph, wherein the knowledge graph includes nodes corresponding to the most unique text elements, and the nodes are interconnected by edges that represent the relationships between the most unique text elements. The method according to claim 1, further comprising:

18. A computer program that ranks multiple text elements, Computer code, wherein the computer code enables one or more sets of processors to perform the following operations, namely: The process involves calculating an embedding vector that locates a first text element among multiple text elements that will be ranked in the embedding space via a word embedding scheme, and An operation to select a set of text fragments from a reference text, wherein each text fragment includes at least one other text element that will be ranked with the first text element, For each text fragment, the operation of calculating the respective distance between the first text element to be ranked and at least one other text element to be ranked, which is located in the embedding space by an embedding vector calculated via the word embedding scheme in the embedding space, The operation involves calculating a specificity score for each text element to be ranked, depending on the respective distance in the embedding space, An operation for storing the aforementioned singularity scores, wherein the singularity scores for the plurality of text elements define the ranking of the text elements by singularity, and Computer code containing instructions and data that cause an operation including A computer program that includes the following features.

19. The computer program according to claim 18, wherein the operation further includes the operation of ranking the plurality of text elements, processing the text data structure using the specificity scores for the text elements, and extracting data having a desired specificity from the text data structure.

20. A computing system for ranking multiple text elements, A set of one or more processors, Machine-readable memory devices and, Computer code stored on the machine-readable storage device, wherein the computer code performs the following operations on the set of one or more processors, namely: The process involves calculating an embedding vector that locates a first text element among multiple text elements that will be ranked in the embedding space via a word embedding scheme, and An operation to select a set of text fragments from a reference text, wherein each text fragment includes at least one other text element that will be ranked with the first text element, For each text fragment, the operation of calculating the respective distance between the first text element to be ranked and at least one other text element to be ranked, which is located in the embedding space by an embedding vector calculated via the word embedding scheme in the embedding space, The operation involves calculating a specificity score for each text element to be ranked, depending on the respective distance in the embedding space, An operation for storing the aforementioned singularity scores, wherein the singularity scores for the plurality of text elements define the ranking of the text elements by singularity, and Computer code having instructions and data that cause an operation including A computing system equipped with [the following features].

21. The computing system according to claim 20, wherein the operation further includes the operation of ranking the plurality of text elements, processing the text data structure using the specificity scores for the text elements, and extracting data having a desired specificity from the text data structure.