suggested query terms

By retrieving nodes of seed words and candidate words from a semi-structured corpus and using the positional relationships of the nodes to determine whether candidate words contain detailed information about seed words, the problem of inaccurate query suggestions under rapidly changing enterprise content data is solved, resulting in more accurate query results.

CN115809320BActive Publication Date: 2026-05-15INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2022-08-26
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

With the rapid changes in enterprise content data, existing technologies lack sufficient word relationship information in search logs, leading to inaccurate query suggestions. For example, "car insurance" is too common for users of "windshield" and cannot effectively improve search results.

Method used

By retrieving nodes of seed words and candidate words from a semi-structured corpus, the positional relationship between nodes is used to determine whether candidate words contain information details about seed words. If confirmed, candidate words are suggested as additional query terms to enhance the query response and retrieval results.

Benefits of technology

It improves the accuracy of query suggestions, enhances the relevance of query results and user experience, and meets the query needs of rapidly changing enterprise content data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115809320B_ABST
    Figure CN115809320B_ABST
Patent Text Reader

Abstract

The present disclosure relates to suggesting query terms. A computer-implemented method for suggesting query terms is provided. The method includes obtaining a seed term and a candidate term. The method also includes retrieving two nodes from a semi-structured corpus that respectively indicate the seed term and the candidate term. The method further includes determining, by a processor device, whether the candidate term includes details of information about the seed term based on a positional relationship between the two nodes in the semi-structured corpus. The method additionally includes suggesting the candidate term as an additional query term in response to a positive determination. The method also includes performing a query with the candidate term as an additional query term in response to a use suggestion acceptance to augment query answer retrieval results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates generally to queries, and more specifically to suggested query terms. Background Technology

[0002] Query suggestions help search engine users improve their search queries. A typical implementation of this feature uses relationships between terms that appear together in search log records or their content data. However, in enterprise environments where content data is used, it changes so rapidly that search logs often lack sufficient information about term relationships. Using secondary relevance from content data also has the drawback of not providing search order information. For example, suggesting "car insurance" for the query "windshield" is too common for users searching for "windshield." Therefore, an improved method for suggesting query terms is needed. Summary of the Invention

[0003] According to various aspects of the present invention, a computer-implemented method for suggesting query terms is provided. The method includes obtaining a seed term and candidate terms. The method further includes retrieving two nodes from a semi-structured corpus that respectively indicate the seed term and the candidate term. The method also includes determining, by a processor device, whether the candidate term includes details of information about the seed term based on the positional relationship between the two nodes in the semi-structured corpus. The method further includes suggesting the candidate term as an additional query term in response to an affirmative determination. The method also includes executing a query with the candidate term as an additional query term in response to an acceptance of the suggestion, to enhance the query response retrieval results.

[0004] According to another aspect of the invention, a computer program product for suggesting query terms is provided. The computer program product includes a non-transitory computer-readable storage medium having program instructions implemented therewith. The program instructions are executable by a computer to cause the computer to perform a method. The method includes obtaining a seed word and candidate words by a processor device. The method further includes retrieving two nodes from a semi-structured corpus, respectively indicating the seed word and the candidate word, by the processor device. The method also includes determining, by the processor device, whether the candidate word includes details of information about the seed word based on the positional relationship between the two nodes in the semi-structured corpus. The method further includes suggesting the candidate word as an additional query term by the processor device in response to an affirmative determination. The method also includes executing a query by the processor device with the candidate word as an additional query term in response to an acceptance of the suggestion, to enhance the query response retrieval results.

[0005] According to other aspects, a computer processing system for suggesting query terms is provided. The computer processing system includes a memory device for storing program code. The computer processor system also includes a processor device operatively coupled to the memory device, the memory device being used to store the program code to obtain seed words and candidate words. The processor device further runs the program code to retrieve two nodes from a semi-structured corpus that respectively indicate the seed word and the candidate word. The processor device further runs the program code to determine, based on the positional relationship between the two nodes in the semi-structured corpus, whether the candidate word includes details of information about the seed word. The processor device further runs the program code to suggest the candidate word as an additional query term in response to an affirmative determination. The processor device further runs the program code to execute the query with the candidate word as an additional query term in response to a suggestion acceptance, thereby enhancing the query response retrieval results.

[0006] These and other features and advantages will become apparent from the following detailed description of illustrative embodiments thereof, which will be read in conjunction with the accompanying drawings. Attached Figure Description

[0007] The following description will provide details of preferred embodiments with reference to the following figures, in which:

[0008] Figure 1 This is a block diagram illustrating an exemplary computing device according to an embodiment of the present invention;

[0009] Figure 2-3 This is a flowchart illustrating an exemplary method according to an embodiment of the present invention;

[0010] Figure 4 This is a diagram illustrating an exemplary portion of a corpus tree according to an embodiment of the present invention;

[0011] Figure 5 This is a diagram illustrating an exemplary corpus tree according to an embodiment of the present invention;

[0012] Figure 6 This is a diagram illustrating an exemplary portion of a corpus tree according to an embodiment of the present invention;

[0013] Figure 7 This is a diagram illustrating an exemplary portion of a corpus tree according to an embodiment of the present invention;

[0014] Figure 8 This is a diagram illustrating an exemplary portion of HTML code according to an embodiment of the present invention;

[0015] Figure 9 This is a diagram illustrating an exemplary search log, a semi-structured corpus, features of positive samples, and features of negative samples according to an embodiment of the present invention;

[0016] Figure 10 This is a block diagram illustrating an illustrative cloud computing environment according to an embodiment of the present invention, having one or more cloud computing nodes communicating with a local computing device used by a cloud consumer; and

[0017] Figure 11 This is a block diagram illustrating a set of functional abstraction layers provided by a cloud computing environment according to an embodiment of the present invention. Detailed Implementation

[0018] The embodiments of the present invention are for suggested query terms.

[0019] Embodiments of the present invention use a semi-structured corpus to restrict the search order of two words based on the relationship between corpus nodes that include each of the two words. In a simple example, the relationship between one node and the other node can be a child node relationship.

[0020] As an example, given a word pair A and B (seed and suggestion), the invention may relate to the following:

[0021] (1) Retrieve nodes N from the corpus that contain A and B respectively. A and N B , where N X = {Nodes including word X}

[0022] (2) Load the configuration of the location constraint R into these two nodes N. A and N B superior.

[0023] Position constraint R (n) A ,n B )∈N A xN B The score mapped to it is ∈ [0,1].

[0024] The score represents node n b Including n a The possibility of details in the information.

[0025] (3) Determine the confidence level of B based on the score.

[0026] Each corpus is organized in a tree structure. These trees are constructed using corpus nodes. A corpus node is a specific point within the corpus structure, and it is used to construct its general structure. Typically, corpus nodes are grouped together based on factors such as geographic location, discourse type, speaker's gender or age, speaker's dialect, target / source language, etc. Here, the present invention utilizes location information within the corpus to suggest query terms.

[0027] Figure 1This is a block diagram illustrating an exemplary computing device 100 according to an embodiment of the present invention. The computing device 100 is configured to suggest search terms.

[0028] The computing device 100 can be implemented as any type of computing or computer device capable of performing the functions described herein, including but not limited to computers, servers, rack-based servers, blade servers, workstations, desktop computers, laptop computers, notebook computers, tablet computers, mobile computing devices, wearable computing devices, network appliances, web appliances, distributed computing systems, processor-based systems, and / or consumer electronics devices. Alternatively or additionally, the computing device 100 can be implemented as one or more computing racks, memory racks, or other racks, chassis, or other components of a physically separate computing device. Figure 1 As shown, computing device 100 illustratively includes processor 110, input / output subsystem 120, memory 130, data storage device 140, and communication subsystem 150, and / or other components and devices typically found in servers or similar computing devices. Of course, in other embodiments, computing device 100 may include other or additional components, such as those typically found in server computers (e.g., various input / output devices). Additionally, in some embodiments, one or more of the illustrative components may be incorporated into another component or otherwise formed part of another component. For example, in some embodiments, memory 130 or a portion thereof may be incorporated into processor 110.

[0029] Processor 110 may be implemented as any type of processor capable of performing the functions described herein. Processor 110 may be implemented as a single processor, multiple processors, a central processing unit (CPU), a graphics processing unit (GPU), a single-core or multi-core processor, a digital signal processor, a microcontroller, or other processor or processing / control circuitry.

[0030] Memory 130 can be implemented as any type of volatile or non-volatile memory or data storage device capable of performing the functions described herein. In operation, memory 130 can store various data and software used during the operation of computing device 100, such as operating systems, applications, programs, libraries, and drivers. Memory 130 is communicatively coupled to processor 110 via I / O subsystem 120, which can be implemented as circuitry and / or components facilitating input / output operations with processor 110, memory 130, and other components of computing device 100. For example, I / O subsystem 120 can be implemented as or otherwise include a memory controller hub, input / output control hub, platform controller hub, integrated control circuitry, firmware devices, communication links (e.g., point-to-point links, bus links, wires, cables, light guides, printed circuit board traces, etc.) and / or other components and subsystems to facilitate input / output operations. In some embodiments, the I / O subsystem 120 may form part of a system-on-a-chip (SOC) and may be integrated on a single integrated circuit chip along with the processor 110, memory 130 and other components of the computing device 100.

[0031] Data storage device 140 can be implemented as one or more devices of any type configured for short-term or long-term data storage, such as, for example, memory devices and circuitry, memory cards, hard disk drives, solid-state drives, or other data storage devices. Data storage device 140 can store program code used to suggest query terms. The communication subsystem 150 of computing device 100 can be implemented as any network interface controller or other communication circuitry, device, or combination thereof capable of enabling communication between computing device 100 and other remote devices via a network. Communication subsystem 150 can be configured to use any one or more communication technologies (e.g., wired or wireless communication) and associated protocols (e.g., Ethernet, etc.). This communication can be achieved using technologies such as WiMAX.

[0032] As shown in the figure, the computing device 100 may also include one or more peripheral devices 160. Peripheral devices 160 may include any number of additional input / output devices, interface devices, and / or other peripheral devices. For example, in some embodiments, peripheral devices 160 may include a display, touchscreen, graphics circuitry, keyboard, mouse, speaker system, microphone, network interface, and / or other input / output devices, interface devices, and / or peripheral devices.

[0033] Of course, the computing device 100 may also include other elements (not shown) that are readily apparent to those skilled in the art, and some elements may be omitted. For example, different other input and / or output devices may be included in the computing device 100, depending on the specific implementation of the computing device 100, as is readily understood by those skilled in the art. For example, different types of wireless and / or wired input and / or output devices may be used. Furthermore, additional processors, controllers, memories, etc., in different configurations may be utilized. Further, in another embodiment, a cloud configuration may be used (e.g., see...). Figure 10-11 Given the teachings of the invention provided herein, those skilled in the art will readily conceive of these and other variations of the processing system 100.

[0034] As used herein, the terms "hardware processor subsystem" or "hardware processor" can refer to a processor, memory (including RAM, cache, etc.), software (including memory management software), or a combination thereof that cooperate to perform one or more specific tasks. In useful embodiments, a hardware processor subsystem may include one or more data processing elements (e.g., logic circuitry, processing circuitry, instruction execution devices, etc.). These one or more data processing elements may be included in a central processing unit, a graphics processing unit, and / or a separate processor- or computing element-based controller (e.g., logic gates, etc.). A hardware processor subsystem may include one or more on-board memories (e.g., cache, dedicated memory array, read-only memory, etc.). In some embodiments, a hardware processor subsystem may include one or more memories (e.g., ROM, RAM, basic input / output system (BIOS), etc.) that may be on-board or off-board, or may be dedicated to use by the hardware processor subsystem.

[0035] In some embodiments, the hardware processor subsystem may include and execute one or more software elements. These one or more software elements may include an operating system and / or one or more applications and / or specific code for achieving a specified result.

[0036] In other embodiments, the hardware processor subsystem may include dedicated, specialized circuitry that performs one or more electronic processing functions to achieve a specified result. Such circuitry may include one or more application-specific integrated circuits (ASICs), FPGAs, and / or PLAs.

[0037] These and other variations of the hardware processor subsystem are also conceived according to embodiments of the present invention.

[0038] Figure 2-3 This is a flowchart illustrating an exemplary method 200 according to an embodiment of the present invention.

[0039] In box 210, obtain the seed word and candidate words.

[0040] In box 220, retrieve two nodes from the semi-structured corpus that indicate the seed word and the candidate word, respectively.

[0041] In box 230, the positional relationship between the two nodes in the semi-structured corpus is used to determine whether the candidate word includes details about the seed word.

[0042] In an embodiment, block 230 includes at least one or more of blocks 230A to 230D.

[0043] In box 230A, predetermined positional constraints are determined for the two nodes, namely, the candidate words need to satisfy details including information about the seed words.

[0044] In box 230B, a trained model is used, which takes as input the shortest path from the seed word to the candidate node and outputs a score indicating the degree to which the candidate word includes details about the seed word.

[0045] In box 230C, the shortest path between these two nodes is configured with a set of directional conditions based on the semi-structured corpus. This condition is used to determine whether a candidate word includes details about the seed word.

[0046] In box 230D, the shortest path between these two nodes is configured as a condition for a set of pairs of node types and directions in the semi-structured corpus. This condition is used to determine whether a candidate word includes details about the seed word.

[0047] In box 240, in response to a positive confirmation, the candidate term is suggested as an additional query term.

[0048] In an embodiment, box 240 may include box 240A.

[0049] In box 240A, a candidate word is preferentially suggested when it appears in a large number of nodes in a semi-structured corpus (e.g., more than 5%, although other percentages may be used depending on the domain) and when it appears in the upper layers of the semi-structured corpus. As an example, an upper layer can be considered as above a half-way marker in the semi-structured corpus. This prioritization can be applied when multiple candidate words are suggested.

[0050] In box 250, in response to the suggestion to accept, the query is executed using the candidate terms as additional query terms.

[0051] Figure 4 This is a diagram illustrating an exemplary portion 400 of a corpus tree according to an embodiment of the present invention.

[0052] Part 400 includes a first node 410 corresponding to the insurance menu and a second node 420 corresponding to auto insurance. The standard for the relationship is child nodes.

[0053] The following example illustrates an unsatisfactory relationship between these two nodes:

[0054] Windshield → Car insurance (bad);

[0055] Replace the windshield (it's faulty);

[0056] The following example illustrates a satisfactory relationship (child nodes) between these two nodes:

[0057] Windshield → Crack (Good).

[0058] Figure 5 This is a diagram illustrating an exemplary corpus tree 500 according to an embodiment of the present invention.

[0059] Circle A represents "windshield". Circle B represents "replacement". The surrounding circle C represents "explosion". In Tree500, "div" represents the parent node and "li" represents the child node.

[0060] R(n A ,n C ) = 0

[0061] R(n A ,n B ) = 1

[0062] It should be understood that R can be configured to disable backward suggestions.

[0063] R should be configured such that if n B It is considered to include n A The more detailed the information in n, the higher R becomes. For example, n A It can be descriptive text, and n B It can be content.

[0064] A description of a simplified implementation of an embodiment of the present invention will now be given.

[0065] In the simplified implementation, n B is n A When the offspring are , R = 1.

[0066] This method works for some Extensible Markup Language (XML) files, but it is not fully capable of representing the structure of Hypertext Markup Language (HTML).

[0067] A description of a flexible implementation version according to embodiments of the present invention will now be given.

[0068] In the flexible implementation version, R uses the data from n A to n B Configure it based on the shortest path limit.

[0069] This approach is applicable to cases where semantic dependencies are not represented by ancestor-descendant pairs.

[0070] Below is an example definition of R using the shortest path:

[0071] If B is within the path that allows it to leave A by moving, then R = 1

[0072] (Path 1) up to the nearest ancestor of li / p

[0073] (Path 2) forwards to at most one sibling of the li / p, or

[0074] (Path 3) Down to descendants at a depth of 10,

[0075] Otherwise, R = 0.

[0076] These three paths are Figure 5 It is illustrated in the text.

[0077] Figure 6 This is a diagram illustrating an exemplary portion 600 of a corpus tree according to an embodiment of the present invention. Portion 600 illustratively illustrates path 1 601, path 2 602, and path 3 603.

[0078] Figure 7 This is a diagram illustrating an exemplary portion 700 of a corpus tree according to an embodiment of the present invention.

[0079] Part 700 includes tire node 710 and windshield node 720. Tire node 710 includes sub-nodes "explosion" and "theft". Windshield node 720 includes sub-nodes "crack", "broken", "replaced", and "stone".

[0080] Visually, in part 700, the "replaced" one looks like the "windshield" of 720.

[0081] Figure 8 This is a diagram illustrating an exemplary portion 800 of HTML code according to an embodiment of the present invention.

[0082] and Figure 7 Compared to part 700 of the corpus tree, part 800 of the HTML code shows that "replaced" is not a child node of "windshield".

[0083] A description of the learning R according to an embodiment of the present invention will now be given.

[0084] R can be automatically configured by learning the shortest paths between nodes, including words in search log records, as positive samples of positional relationships.

[0085] For each record in the search log:

[0086] (1) Obtain nodes from a semi-structured corpus that includes every word in the record. For example, nodes that include "windshield" and nodes that include "replaced".

[0087] (2) For each pair of nodes: node A containing the word and node B containing another word:

[0088] (2A) Create features in the form of (node ​​type, direction) based on the nodes on the shortest path from A to B; and

[0089] (2B) Features of negative samples are created by adding a node outside the path to the subpath. For example, an additional "upward" move or an additional "generational move" can be added.

[0090] Figure 9 This is a diagram illustrating an exemplary search log 910, a semi-structured corpus 920, features 930 of positive samples, and features 940 of negative samples according to an embodiment of the present invention.

[0091] Search log 910 includes nodes with the words "car", "lid", and "bicycle" on one line / level and "windshield" and "replaced" on the next line / level.

[0092] In the semi-structured corpus 920, circle "A" represents "windshield" and circle "B" represents "replaced".

[0093] Regarding node type 920 in the semi-structured corpus:

[0094] "div" represents the "div" tag that represents a box area in an HTML file;

[0095] "h2" represents the "h2" tag, which signifies a second-level heading in an HTML file;

[0096] "ul" represents the "ul" tag that represents an unordered list in an HTML file;

[0097] "li" represents the "li" tag that indicates an item in an unordered list within an HTML file; and

[0098] "p" represents the "p" tag in an HTML file.

[0099] Based on the semi-structured corpus 920, the characteristics of positive samples include:

[0100] (h2, upward), (ul, same generation), (li, downward), (p, downward).

[0101] Based on the semi-structured corpus 920, the characteristics of negative samples include:

[0102] (h2, up), (div, up)

[0103] (h2, upward), (ul, sibling), (div, sibling).

[0104] Depending on the implementation method, the following options can be used:

[0105] (1) Features may include counts. For example, (div, up, 1), (div, up, 2).

[0106] (2) Nodes in the sub-path can be excluded from negative samples. For example, (h2, up), (div, up).

[0107] In one embodiment, Figure 2 One or more boxes of method 200 can be executed in the cloud.

[0108] It should be understood that while this disclosure includes a detailed description of cloud computing, the implementation of the teachings cited herein is not limited to cloud computing environments. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed hereafter.

[0109] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model may include at least five features, at least three service models, and at least four deployment models.

[0110] The features are as follows:

[0111] On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power, such as server time and network storage, as needed, without requiring human interaction with the service provider.

[0112] Extensive network access: Capabilities are available through networks and accessed via standard mechanisms that facilitate the use of heterogeneous thin client or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0113] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned as needed. There is a sense of location independence because consumers typically do not have control or knowledge of the exact location of the resources provided, but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0114] Rapid flexibility: The ability to provide capacity quickly and flexibly, automatically scaling down and up rapidly in some situations to scale up rapidly. For consumers, the available supply capacity often appears unlimited and can be purchased in any quantity at any time.

[0115] Measuring services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.

[0116] The service model is as follows:

[0117] Software as a Service (SaaS): This provides consumers with the ability to use the provider's applications running on cloud infrastructure. Applications can be accessed from different client devices via thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application capabilities, with possible exceptions such as limited user-specific application configuration settings.

[0118] Platform as a Service (PaaS): This provides consumers with the ability to deploy applications created or acquired by the consumer using programming languages ​​and tools supported by the provider onto cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and the configuration of any application hosting environment.

[0119] Infrastructure as a Service (IaaS): The capabilities provided to consumers are processing, storage, networking, and other basic computing resources that enable consumers to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but rather have control over the operating system, storage, deployed applications, and potentially limited control over selected networking components (e.g., host firewalls).

[0120] The deployment model is as follows:

[0121] Private cloud: A cloud infrastructure that operates solely for an organization. It can be managed by the organization or a third party and can exist on-site or off-site.

[0122] Community cloud: A cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.

[0123] Public cloud: Makes cloud infrastructure available to the public or large industry groups and is owned by an organization that sells cloud services.

[0124] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and applications to be ported (e.g., cloud bursting for load balancing between clouds).

[0125] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure comprising a network of interconnected nodes.

[0126] See now Figure 10 The diagram illustrates an illustrative cloud computing environment 1050. As shown, the cloud computing environment 1050 includes one or more cloud computing nodes 1010 that can communicate with local computing devices used by cloud consumers, such as, for example, personal digital assistants (PDAs) or cellular phones 1054A, desktop computers 1054B, laptop computers 1054C, and / or automotive computer systems 1054N. The nodes 1010 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment 1050 to provide infrastructure, platforms, and / or software as services that cloud consumers do not need to maintain on their local computing devices. It should be understood that... Figure 10 The types of computing devices 1054A-N shown are intended to be illustrative only, and computing node 1010 and cloud computing environment 1050 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).

[0127] See now Figure 11 This demonstrates the 1050 cloud computing environment ( Figure 10 This provides a set of functional abstractions. It should be understood beforehand that... Figure 11 The components, layers, and functions shown are intended to be illustrative only, and embodiments of the invention are not limited thereto. As described, the following layers and corresponding functions are provided:

[0128] The hardware and software layer 1160 includes hardware and software components. Examples of hardware components include: a host 1161; a server 1162 based on a RISC (Reduced Instruction Set Computer) architecture; a server 1163; a blade server 1164; a storage device 1165; and a network and network components 1166. In some embodiments, the software components include network application server software 1167 and database software 1168.

[0129] The virtualization layer 1170 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 1171; virtual storage 1172; virtual network 1173, including virtual private network; virtual application and operating system 1174; and virtual client 1175.

[0130] In one example, management layer 1180 can provide the following functionalities: Resource Provisioning 1181 Provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and Pricing 1182 Provides cost tracking as resources are utilized within the cloud computing environment and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security Provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User Portal 1183 Provides consumers and system administrators with access to the cloud computing environment. Service Level Management 1184 Provides cloud resource allocation and management to ensure that required service levels are met. Service Level Agreement (SLA) Planning and Fulfillment 1185 Provides pre-scheduling and procurement of cloud resources, anticipating future requirements for those resources according to the SLA.

[0131] The workload layer 1190 provides examples of functionalities that can be leveraged in a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 1191; software development and lifecycle management 1192; virtual classroom education delivery 1193; data analysis and processing 1194; transaction processing 1195; and suggested query terms 1196.

[0132] This invention can be a system, method, and / or computer program product with any possible level of technical detail integration. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to execute aspects of the invention.

[0133] Computer-readable storage media can be tangible means for retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital universal disk (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or protrusions in slots having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.

[0134] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.

[0135] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as SMALLTALK, C++, etc.) and conventional procedural programming languages ​​(such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuitry in order to perform aspects of this invention.

[0136] The present invention will now be described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0137] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.

[0138] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce computer-implemented processing, such that the instructions executed on the computer, other programmable apparatus, or other device perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than indicated in the figures. For example, depending on the functions involved, two consecutively shown blocks may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.

[0140] References to the invention in this specification as "one embodiment" or "embodiment" and other variations thereof mean that a particular feature, structure, characteristic, etc., described in connection with that embodiment is included in at least one embodiment of the invention. Therefore, the phrases "in one embodiment" or "in an embodiment" appearing in various places throughout the specification, as well as any other variations, do not necessarily refer to the same embodiment.

[0141] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one of” is intended to include selecting only the first listed item (A), or only the second listed item (B), or selecting both options (A and B). As another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” this wording is intended to cover only the selection of the first listed option (A), or only the selection of the second listed option (B), or only the selection of the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed items (A and C), or only the second and third listed items (B and C), or all three options (A, B, and C). It will be apparent to those skilled in the art that this can be extended for many of the listed items.

[0142] Preferred embodiments of the systems and methods have been described (these are intended to be illustrative and not restrictive), and it should be noted that modifications and variations can be made by those skilled in the art based on the foregoing teachings. Therefore, it should be understood that changes may be made to the specific embodiments disclosed within the scope of the invention as outlined in the appended claims. Various aspects of the invention, having the details and features required by patent law, have thus been described, and the claimed and desired protection by a patent certificate is set forth in the claims.

Claims

1. A computer-implemented method for suggesting query terms, comprising: Obtain seed words and candidate words; Retrieve two nodes that respectively indicate the seed word and the candidate word from a semi-structured corpus, the semi-structured corpus having a hierarchical structure of nodes grouped based on position information, the position information including the relationship between nodes based on the syntax of the semi-structured corpus; The processor determines whether the candidate word includes details of information about the seed word based on positional relationships, where the positional relationships include the order of the two nodes in the semi-structured corpus based on the positional information. In response to a positive confirmation, the aforementioned candidate terms are suggested as additional query terms; as well as In response to the user's acceptance of the usage suggestion, the query is executed using the candidate terms as additional query terms to enhance the query response retrieval results.

2. The computer-implemented method according to claim 1, wherein, The determination step includes determining predetermined positional constraints on the two nodes, namely, the candidate word needs to satisfy details including information about the seed word.

3. The computer-implemented method according to claim 2, wherein, The determination step includes using a trained model that takes as input the shortest path from the seed word to the candidate node and outputs a score indicating the degree to which the candidate word includes details about the seed word.

4. The computer-implemented method according to claim 1, wherein, The seed words are included in queries that add the candidate words in response to the user accepting the usage suggestion.

5. The computer-implemented method according to claim 1, wherein, When the candidate word appears in a descendant node of a given node among the two nodes that include the query word, the candidate word is determined to include details about the seed word.

6. The computer-implemented method according to claim 1, wherein, The shortest path between the two nodes is configured with conditions for a set of node types in the semi-structured corpus, which are used to determine whether the candidate word includes details about the seed word.

7. The computer-implemented method according to claim 1, wherein, The shortest path between the two nodes is configured with a set of directional conditions in the semi-structured corpus, which are used to determine whether the candidate word includes details about the seed word.

8. The computer-implemented method according to claim 1, wherein, The shortest path between the two nodes is configured as a condition for a set of pairs of node types and directions in the semi-structured corpus, the condition being used to determine whether the candidate word includes details about the seed word.

9. The computer-implemented method according to claim 1, wherein, When a candidate word appears in more than 5% of multiple nodes in the semi-structured corpus and appears in the upper layer of the semi-structured corpus, the candidate word is given priority for suggestion, the upper layer being above the halfway marker in the semi-structured corpus.

10. The computer-implemented method according to claim 1, wherein, The parameters for prioritizing suggestions for the candidate words over other candidate words with non-shortest paths to the seed word are learned by learning the shortest paths between the seed word and the candidate words in the semi-structured corpus.

11. The computer-implemented method according to claim 10, wherein, Using positive and negative samples for learning, the positive samples have the shortest path, and the negative samples are modified positive samples that are added to existing paths to extend the existing paths to the non-shortest paths.

12. A computer program product for suggesting query terms, the computer program product comprising program instructions executable by a computer to cause the computer to perform the method according to any one of claims 1-11.

13. A computer processing system for suggesting query terms, comprising: Storage device for storing program code; as well as A processor device, operatively coupled to the storage device, for storing program code for performing the method according to any one of claims 1-11.