Information retrieval method and device, electronic equipment and program product

By storing and retrieving text through a cascading tree structure, the problems of low retrieval efficiency and difficulty in information extraction in existing technologies are solved, and efficient and accurate information retrieval is achieved, which is suitable for complex scenarios such as corporate meeting texts.

CN120687586APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510244831.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

When the text content is complex and the volume is huge, existing technologies find it difficult to retrieve effective information efficiently and accurately. In particular, when processing large amounts of text, the retrieval efficiency is low, key information extraction and organization are difficult, and it is difficult to quickly locate specific information.

Method used

A cascading tree structure is used to store text, and a multi-level structure is formed by node clustering. Nodes matching the query are searched layer by layer until the text block at the last level is found, realizing information retrieval from abstract to concrete, avoiding searching and analyzing each sentence one by one.

Benefits of technology

It significantly improves retrieval efficiency and accuracy, and can quickly and accurately extract required information from text to meet real-time application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687586A_ABST
    Figure CN120687586A_ABST
Patent Text Reader

Abstract

The invention discloses an information retrieval method and device, electronic equipment and a program product, which are used for efficiently and accurately retrieving useful information in a text. The method comprises the following steps: in response to a first query request, determining a first node matched with the first query request from nodes of a first level of a cascade tree; the cascade tree comprises a plurality of nodes, the plurality of nodes are distributed in a plurality of hierarchies, the nodes in the previous hierarchy are obtained by clustering the nodes in the next hierarchy, the nodes in the last hierarchy correspond to text blocks of the first text, and the nodes in the non-last hierarchy correspond to themes of the text blocks; determining a second node matched with the first query request from the next level of the level where the first node is located; and if the determined second node is the node in the last hierarchy, determining a query result corresponding to the first query request based on the text block corresponding to the second node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information processing technology, and in particular to an information retrieval method, device, electronic device, and program product. Background Art

[0002] In some business scenarios, it is necessary to retrieve effective and accurate information from text. However, when the text content is complex and the amount of text is huge, it is often difficult to retrieve useful information from the text efficiently and accurately. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide an information retrieval method, device, electronic device and program product for efficiently and accurately retrieving useful information from a text.

[0004] In order to achieve the above objectives, the embodiments of the present application adopt the following technical solutions: In a first aspect, an embodiment of the present application provides an information retrieval method, comprising: In response to a first query request, determining a first node that matches the first query request from nodes in a first level of a cascade tree; the cascade tree includes a plurality of nodes, the plurality of nodes being distributed in a plurality of levels, the nodes in a previous level being obtained by clustering nodes in a next lower level, the nodes in the last level corresponding to a text block of the first text, and the nodes in a non-last level corresponding to a topic of the text block; Determining, from a level below the level where the first node is located, a second node that matches the first query request; If the determined second node is a node in the last level, a query result corresponding to the first query request is determined based on the text block corresponding to the second node.

[0005] In a second aspect, an embodiment of the present application provides an information retrieval device, comprising: a determination module configured to, in response to receiving a first query request, determine a first node that matches the first query request from nodes in a first level of a cascade tree; the cascade tree comprising a plurality of nodes, the plurality of nodes being distributed in a plurality of levels, the nodes in a previous level being obtained by clustering nodes in a next lower level, the nodes in a final level corresponding to a text block of the first text, and the nodes in a non-final level corresponding to a topic of the text block; The determining module is further configured to determine a second node matching the first query request from nodes in a layer below the layer where the first node is located; The query module is configured to determine a query result corresponding to the first query request based on a text block corresponding to the second node if the determined second node is a node in the last level.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, including: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the information retrieval method provided in the first aspect.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the information retrieval method provided in the first aspect.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute some or all of the steps in the information retrieval method provided in the first aspect.

[0009] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: The text is stored in a cascading tree structure. The cascading tree includes multiple nodes, which are distributed in multiple levels. The nodes in the previous level are obtained by clustering the nodes in the next level. The nodes in the last level correspond to the text blocks of the text, and the nodes in the non-last level correspond to the topics of the text blocks. In this way, the text is organized and stored in a cascading tree structure. Under this structure, the content of the text is gradually refined into individual text blocks according to the most abstract topics. On this basis, when retrieving the text, starting from the nodes in the first level of the cascade tree, the nodes that match the query request are searched layer by layer until the matching nodes in the last level are obtained. That is, the text blocks that match the query request are found from the text in a way from abstract to concrete topics, without having to search and analyze each sentence in each text one by one. This not only significantly improves the retrieval efficiency, but also ensures that the found text blocks contain the required information, and then the required information can be quickly and accurately extracted from the text blocks. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A flowchart of an information retrieval method provided in accordance with an embodiment of the present application; Figure 2 A schematic diagram of a cascade tree structure provided for one embodiment of the present application; Figure 3A schematic diagram of a flow chart of a method for constructing a cascade tree provided in one embodiment of the present application; Figure 4 A schematic diagram of determining a first sentence provided in accordance with an embodiment of the present application; Figure 5 A schematic diagram of a cascade tree construction process provided in one embodiment of the present application; Figure 6 A schematic structural diagram of an information retrieval device provided in one embodiment of the present application; Figure 7 A schematic structural diagram of an electronic device provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0011] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0012] The terms "first," "second," and the like in this specification and claims are used to distinguish similar objects and are not intended to describe a particular order or precedence. It should be understood that such terms are interchangeable where appropriate so that the embodiments of the present application can be implemented in sequences other than those illustrated or described herein. In addition, the term "and / or" in this specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the connected objects are in an "or" relationship.

[0013] As mentioned in the background, in some business scenarios, it is necessary to retrieve valid and accurate information from text. For example, some companies may hold hundreds or even thousands of meetings every day, and the text of each meeting is complex and large in volume. Retrieving valid and accurate information from such a large amount of meeting text is often difficult.

[0014] In related technologies, the Retrieval-Augmented Generation (RAG) method is usually used for retrieval. That is, relevant information is retrieved from an external knowledge base and input into large-scale language models (LLMs) as prompts to enhance the ability of large-scale language models to handle knowledge-intensive tasks and obtain key information in the text.

[0015] However, this retrieval method often has the following drawbacks when faced with large amounts of text: (1) Low retrieval efficiency. When processing large amounts of text, such as texts from multiple long meetings, the RAG method needs to retrieve and analyze each sentence in each text one by one, resulting in low processing efficiency and unable to meet the needs of real-time applications.

[0016] (2) Difficulty in extracting and organizing key information. The above-mentioned RAG methods usually rely on simple text similarity or keyword matching, which makes it difficult to effectively extract and organize all key information in the text in complex business scenarios (such as meeting scenarios).

[0017] (3) It is difficult to quickly locate and retrieve specific information. Faced with massive amounts of text, it is often difficult to quickly find the required specific information through simple queries, which limits the practicality of the retrieval system and the user experience.

[0018] In view of this, an embodiment of the present application proposes an information retrieval method, which stores text in a cascade tree structure. The cascade tree includes multiple nodes, which are distributed in multiple levels. The nodes in the upper level are obtained by clustering the nodes in the lower level. The nodes in the last level correspond to the text blocks of the text, and the nodes in the non-last level correspond to the topics of the text blocks. In this way, the text is organized and stored in a cascade tree structure. Under this structure, the content of the text is gradually refined into individual text blocks according to the most abstract topic. On this basis, when searching the text, starting from the nodes of the first level of the cascade tree, the nodes that match the query request are searched layer by layer until the matching nodes in the last level are obtained. That is, the text blocks that match the query request are found from the text in a way from abstract to concrete topics, without having to search and analyze each sentence in each text one by one. This not only significantly improves the retrieval efficiency, but also ensures that the found text blocks contain the required information, and then the required information can be quickly and accurately extracted from the text blocks.

[0019] It should be understood that the information retrieval method provided in the embodiments of the present application can be executed by an electronic device, specifically by a processor of the electronic device. The electronic device herein may include a terminal, such as but not limited to a smartphone, tablet computer, laptop computer, desktop computer, intelligent voice interaction device, smart home appliance, smart watch, vehicle-mounted terminal, aircraft, etc.; or the electronic device may also include a server, such as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0020] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0021] Please refer to Figure 1, is a flow chart of an information retrieval method provided in one embodiment of the present application, the method comprising: S102 : In response to a first query request, determine a first node matching the first query request from nodes in a first level of the cascade tree.

[0022] A cascade tree is a data structure that is often used to represent data with hierarchical relationships, such as organizational structures, departmental relationships, etc. The cascade tree represents the hierarchical relationship between data through nodes and edges. Each node represents an entity, and the edge represents the hierarchy or parent-child relationship between entities. In this application, a cascade tree is used to represent texts with a hierarchical relationship. These texts can be different types of texts belonging to the same department, or they can be multiple different texts under the same topic. For example, a cascade tree is used to store meeting texts, which are further divided into texts of multiple meetings, such as the text of meeting 1, the text of meeting 2, and the text of meeting 3, etc. The meeting text of each meeting can have a cascade tree, and the meeting texts of all meetings can form a larger cascade tree.

[0023] The first query request carries the second text, and is used to request to search the first text for text content that matches the second text.

[0024] The cascade tree includes multiple nodes distributed in multiple levels. The nodes in the previous level are obtained by clustering the nodes in the next level. The nodes in the last level correspond to the text block of the first text, and the nodes in the non-last level correspond to the topic of the text block.

[0025] For example, Figure 2As shown, the first text includes the texts of 4 meetings, and each meeting has a corresponding cascade tree. Taking Meeting 1 as an example, the cascade tree corresponding to Meeting 1 includes multiple nodes, and the multiple nodes are distributed in three levels. The nodes in the third level correspond to the text blocks of the text of Meeting 1; the nodes in the second level are obtained by clustering the nodes in the third level, and the nodes in the second level correspond to the topics of the text blocks corresponding to the nodes in the third level; the nodes in the first level are obtained by clustering the nodes in the second level, and the nodes in the first level represent more abstract topics that are further refined from the topics corresponding to the nodes in the second level, that is, the topic of Meeting 1. For example, the theme corresponding to the nodes in the first level is "Innovation and Exploration of Brand Marketing Concepts", and the themes corresponding to the nodes in each second level are "Discussion on New Product Launch Promotion Plan Using Social Media" and "Planning Meeting for Offline Store Promotion Activities in the Second Quarter of This Year". The text block corresponding to the node in the third level connected to the node in the second level "Discussion on New Product Launch Promotion Plan Using Social Media" includes text related to the theme, and the text block corresponding to the node in the third level connected to the node in the second level "Planning Meeting for Offline Store Promotion Activities in the Second Quarter of This Year" includes text related to the theme.

[0026] Furthermore, considering that different conferences may have similar themes, we can cluster the nodes in the first level of each conference to obtain the nodes in the previous level. This operation is repeated until the nodes in the previous level cannot be clustered, thus obtaining a larger cascade tree. The nodes in the previous level correspond to a more abstract theme, which is a further refinement of the theme corresponding to the nodes in the next level.

[0027] For example, clustering the nodes in the first level of each of Conferences 1 to 4 yields three nodes in the previous level. In the previous level, the first node is connected to the nodes in the first level of Conference 1 and the nodes in the first level of Conference 2, and its corresponding topic is a refinement of the topics of Conferences 1 and 2. The second node is connected to the nodes in the first level of Conference 3, and its corresponding topic is a refinement of the topic of Conference 3. The third node is connected to the nodes in the first level of Conference 4, and its corresponding topic is a refinement of the topic of Conference 4. Clustering these three nodes yields a root node, whose corresponding topic is a refinement of the topics corresponding to each of these three nodes.

[0028] In the implementation of this application, each node in the cascading tree may have a corresponding label, and the node label is used to describe the topic or text block corresponding to the node. As an example, the node label may include a first label and a second label, where the first label is used to describe the summary of the topic or text block corresponding to the node, and the second label is used to describe the key information of the topic or text block corresponding to the node, such as keywords, main content, important decisions, etc.

[0029] In practical applications, the labels of nodes can be dynamically adjusted according to the actual needs of users to ensure that query results matching the first query request are returned, thereby improving the accuracy and efficiency of information retrieval.

[0030] In the embodiment of the present application, the above S102 can be achieved through various appropriate methods.

[0031] In one embodiment, the first query request is embedded to obtain a first vector; the topics corresponding to the nodes in the first level are embedded to obtain a second vector of the nodes in the first level; and from the first level, a node whose similarity between the second vector and the first vector is greater than or equal to a similarity threshold is selected as the first node matching the first query request.

[0032] In another embodiment, the first query request is embedded to obtain a first vector; the labels of the nodes in the first level are embedded to obtain a second vector of the nodes in the first level; and from the first level, a node whose similarity between the second vector and the first vector is greater than or equal to a similarity threshold is selected as the first node that matches the first query request.

[0033] Specifically, since the first label of a node describes the summary of the topic corresponding to the node, in order to improve retrieval efficiency, the first labels of the nodes in the first level can be embedded to obtain the second vectors of the nodes in the first level.

[0034] The above describes some implementation methods of the above S102. Of course, it should be understood that the above S102 can also be implemented in other ways, and the present embodiment of the application does not limit this.

[0035] S104: Determine a second node that matches the first query request from a level below the level where the first node is located.

[0036] In one embodiment, the above S104 includes the following steps: Step A1: Embed the second text to obtain a first vector.

[0037] Step A2: For each node in the next level, embed the label of the node to obtain the second vector of the node.

[0038] Specifically, for each node in the next layer, the first label of the node may be embedded to obtain the second vector of the node.

[0039] Step A3: determining a second node matching the first query request from the nodes in the next layer based on the first vector and the second vector of each node in the next layer.

[0040] As an example, the similarity between the first vector and the second vector of the node in the next layer is determined; from the next layer, a node whose second vector has a similarity with the first vector greater than a similarity threshold is selected as the second node matching the first query request.

[0041] For example, Figure 2 Taking the cascade tree corresponding to meeting 1 as an example, if the similarity between the second vector and the first vector of node 1 in the second level is greater than the similarity threshold, node 1 is determined to be the second node.

[0042] As another example, the label of the first node is embedded to obtain the second vector of the first node; for each node at the next level, the second vector of the node is projected onto the second vector of the first node to obtain the third vector of the node; and the difference between the second vector of the first node and the third vector of the node is determined as the fourth vector of the node; from the nodes at the next level, the node whose similarity between the fourth vector and the first vector is greater than or equal to the similarity threshold is determined as the second node that matches the first query request.

[0043] Specifically, assuming that p is the projection matrix and x is the second vector of the first node, then the fourth vector of the node at the next level is r=x-px.

[0044] For example, continue with Figure 2 Taking the cascade tree corresponding to meeting 1 in the second level as an example, assuming that node 1 in the first level is the first node, the second vector of node 1 in the second level is projected to the second vector of the first node, and the projection vector of the second vector of node 1 in the second level in the direction of the second vector of the first node is obtained. The straight line where this projection vector is located is the direction of the straight line represented by the second vector of the first node, and the projection vector is the third vector of node 1 in the second level; then, the difference between the third vector of node 1 in the second level and the second vector of the first node is determined as the fourth vector of node 1 in the second level, and the fourth vector reflects the information difference between node 1 in the second level and the first node; further, the similarity between the fourth vector of node 1 in the second level and the first vector is determined. If the similarity is greater than or equal to the similarity threshold, node 1 in the second level is determined as the second node.

[0045] By performing similar operations on the node 2 in the second level, it is possible to determine whether the node 2 in the second level is the second node.

[0046] In practical applications, the similarity threshold can be set according to actual needs. For example, the smaller the similarity threshold, the greater the number of second nodes, the higher the recall rate, and the lower the precision rate.

[0047] In the above embodiment, by projecting the second vector of the node in the next level to the second vector of the first node, the difference between the projected vector and the second vector of the first node implies the information difference between the first node and the node in the next level. The second node is screened based on the similarity between the difference and the first vector, which can ensure that the second node is highly matched with the first query request, thereby improving the retrieval accuracy.

[0048] The above describes some implementation methods of the above S104. Of course, it should be understood that the above S104 can also be implemented in other ways, and the present embodiment of the application does not limit this.

[0049] S106: If the determined second node is a node in the last level, a query result corresponding to the first query request is determined based on the text block corresponding to the second node.

[0050] In another embodiment, if the determined second node is not a node in the last level, the second node is determined as the first node, and the above S104 is executed again.

[0051] For example, Figure 2 Taking the cascade tree corresponding to conference 1 as an example, if the first node is node 1 in the first level, then the second node matching the first query request is determined from the second level. If the determined second node is node 1 in the second level, then node 1 is used as the new first node, and the second node matching the first query request is determined from the nodes in the third level connected to the new first node. Since the second node is a node in the last level, the process stops.

[0052] In the above S106 , the query result corresponding to the first query request may be determined in various appropriate ways.

[0053] In one embodiment, if the determined second node is a node in the last level, the text block corresponding to the second node is determined as the query result corresponding to the first query request.

[0054] In another embodiment, since the first label of the node in the last level describes the summary of the text block corresponding to the node in the last level, in order to facilitate the rapid acquisition of the main content of the text block, if the second node determined is a node in the last level, the first label of the second node is determined as the query result corresponding to the first query request.

[0055] In another embodiment, since the second label of the node in the last level describes the key information of the text block corresponding to the node in the last level, in order to facilitate the rapid acquisition of the key content of the text block, if the determined second node is a node in the last level, the second label of the second node is determined as the query result corresponding to the first query request.

[0056] The above describes some implementation methods of the above S106. Of course, it should be understood that the above S106 can also be implemented in other ways, and the present embodiment of the application does not limit this.

[0057] The information retrieval method provided in the embodiment of the present application stores text in a cascade tree structure. The cascade tree includes multiple nodes, which are distributed in multiple levels. The nodes in the upper level are obtained by clustering the nodes in the lower level. The nodes in the last level correspond to the text blocks of the text, and the nodes in the non-last level correspond to the topics of the text blocks. In this way, the text is organized and stored in a cascade tree structure. Under this structure, the content of the text is gradually refined into individual text blocks according to the most abstract topic. On this basis, when searching the text, starting from the nodes of the first level of the cascade tree, the nodes that match the query request are searched downward layer by layer until the matching nodes in the last level are obtained. That is, the text blocks that match the query request are found from the text in a manner from abstract to specific in terms of the topic, without the need to search and analyze each sentence in each text one by one. This not only significantly improves the retrieval efficiency, but also ensures that the found text blocks contain the required information, and then the required information can be quickly and accurately extracted from the text blocks.

[0058] The present application embodiment also provides a method for constructing a cascade tree, which can be implemented before the above S102. Figure 3 , is a flow chart of a method for constructing a cascade tree provided in one embodiment of the present application, the method comprising the following steps: S302: Perform semantic segmentation on the first text to obtain multiple text blocks.

[0059] In one embodiment, since punctuation marks play an important role in dividing semantic units in a text, for example, a period ".", a question mark "?", an exclamation mark "!", etc. usually mark the end of a sentence, and a semicolon ";" can divide sentences that are semantically independent but related, by identifying these punctuation marks, the first text can be divided into sentence-level text blocks.

[0060] For example, for the first text "The weather is very good today. Let's go hiking? We can enjoy the scenery on the way", the first text can be divided into three text blocks based on the question mark and period: "The weather is very good today.", "Let's go hiking?", "We can enjoy the scenery on the way".

[0061] In another embodiment, a part-of-speech tagging tool is used to tag the part of speech of each word in the first text, such as noun, verb, adjective, etc.; then, the first text is subjected to syntactic analysis based on the part of speech of each word, the grammatical structure of the first text is parsed, and the subject, predicate, object and other components in each sentence of the first text are determined. Through this information, semantically closely related phrases or words can be identified, and the first text can be divided into text blocks based on phrases.

[0062] In another embodiment, the above S302 includes the following steps: Step B1: Combine two adjacent sentences in the first text to obtain multiple sentence pairs.

[0063] For example, the first text contains n sentences. For the i-th sentence, the i-th sentence and the i+1-th sentence can be combined to obtain multiple sentence pairs.

[0064] For another example, if the first text is a meeting text, which includes questions and answers during the meeting, the corresponding questions and answers are concatenated into one sentence, and then adjacent sentences in the first text are combined to obtain multiple sentence pairs.

[0065] Step B2: For each sentence pair, determine a first score for the sentence pair based on the similarity between the two sentences in the sentence pair.

[0066] Specifically, for each sentence pair, SBERT is used to convert each sentence in the pair into a high-dimensional vector to capture the semantic information of each sentence; then, the similarity between the high-dimensional vectors of each sentence in the pair is calculated and determined as the first score of the sentence pair.

[0067] Step B3: For each sentence in the first text, determine a second score of the sentence based on the first score of the first sentence pair in which the sentence is located and the first score of the second sentence pair adjacent to the first sentence pair.

[0068] Each sentence has a corresponding first sentence pair. The first sentence pair corresponding to each sentence includes the sentence and the next sentence of the sentence. For example, for the i-th sentence, the first sentence pair corresponding to the sentence includes the i-th sentence and the i+1th sentence , expressed as . Accordingly, the second sentence pair adjacent to the first sentence pair includes and .

[0069] Each sentence has a corresponding second score, which indicates the degree of semantic difference between the sentence and adjacent sentences.

[0070] As an example, the above step B3 includes: obtaining the difference between the first score of the second sentence pair and the first score of the first sentence pair as the second score of the sentence.

[0071] For example, take the i-th sentence For example, the first sentence pair is , the second sentence pair includes and , then the sentence The second score is The first score and After summing the first score, subtract twice the value of the first score.

[0072] As another example, in order to make the second score of a sentence more accurately reflect the degree of semantic difference between the sentence and its adjacent sentences to improve retrieval accuracy, the above-mentioned step B3 includes: obtaining the difference between the first score of the second sentence pair and the first score of the first sentence pair to obtain the third score of the sentence; and obtaining the second score of the sentence based on the average of the third score of the sentence and the third score of the second sentence within the neighborhood of the sentence.

[0073] For example, take the i-th sentence For example, the third score of the sentence can be determined by the following formulas (1) to (3).

[0074] (1) (2) (3) in, Expressing sentences The third score; Indicates the second sentence pair The first score of The monotonically increasing maximum similarity on the left is Figure 4As indicated by the arrow on the left side of the middle circle; Indicates the second sentence pair The first score of The monotonically increasing maximum similarity on the right side is Figure 4 As shown by the arrow on the right side of the middle circle; Indicates the first sentence pair The first score.

[0075] Further, using the sentence The third score of the k adjacent sentences on the left and right, for the sentence The third score is smoothed to obtain the sentence The second score is expressed as By smoothing the third score of a sentence, it helps to eliminate abnormal third scores and provides reliable data support for improving retrieval accuracy.

[0076] Step B4: dividing the first text into a plurality of text blocks based on the first sentence in the first text whose second score is greater than or equal to the score threshold.

[0077] Specifically, the sentences in the first text may be sorted in descending order of the second scores, and the first N sentences may be selected as segmentation points to divide the first text into a plurality of text blocks.

[0078] For example, the first text includes four sentences A, B, C, and D. Sentence B is determined as the segmentation point through the above method, and the first text is divided into two text blocks. One text block includes sentences A and B, recorded as [A, B]; the other text block includes sentences C and D, recorded as [C, D].

[0079] In the above embodiment, since the second score of each sentence reflects the degree of semantic difference between the sentence and the adjacent sentences, the first sentence whose second score is greater than or equal to the score threshold is a semantic turning point in the first text. Based on this, using the first sentence as a segmentation point can accurately segment the first text into multiple text blocks with different semantics, providing data support for the subsequent construction of a cascade tree from abstract to concrete and from macro to micro.

[0080] The above describes some implementation methods of the above S302. Of course, it should be understood that the above S302 can also be implemented in other ways, and the present embodiment of the application does not limit this.

[0081] S304: Create multiple nodes corresponding to the multiple text blocks one by one as nodes in the last level.

[0082] For example, Figure 5As shown, the sentences in the first text are semantically segmented to obtain 5 text blocks, and then 5 nodes are created as nodes in the last level, each node corresponding to a text block.

[0083] S306 : Determine labels of the nodes in the last level based on the text blocks corresponding to the nodes in the last level.

[0084] Specifically, a summary can be generated for the text block corresponding to the node in the last level, and the generated summary can be used as the first label of the node in the last level; key information can be extracted from the text block corresponding to the node in the last level, and the extracted key information can be used as the second label of the node in the last level.

[0085] Generating a summary for a text block can be achieved through various appropriate methods, which are not limited in the present embodiment. For example, for each text block, a first prompt word indicating summary generation is input into a large-scale language model, and the semantic understanding and text generation capabilities of the large-scale language model are utilized to generate a summary for the text block, which is then used as the first label for the node in the last level corresponding to the text block.

[0086] Extracting key information from a text block can be accomplished in various appropriate ways, which are not discussed in the present embodiment. For example, for each text block, a second prompt word indicating key information extraction is input into a large-scale language model, and the semantic understanding and text processing capabilities of the large-scale language model are utilized to extract the key information of the text block, and then the key information is used as the second label of the node in the last level corresponding to the text block.

[0087] S308 , starting from the last level, clustering the nodes in each level based on the labels of the nodes in each level, and obtaining the nodes in the previous level and the labels of the nodes in the previous level.

[0088] In one embodiment, the above S308 includes the following steps: Step C1: for each level, cluster the nodes in the level based on the labels of the nodes in the level to obtain at least one cluster, each cluster including at least one node.

[0089] In the embodiment of the present application, clustering of nodes in each level can be achieved through various clustering algorithms, which is not limited in the embodiment of the present application.

[0090] As an example, since the label of the node includes a first label, the first label describes the summary of the topic or text block corresponding to the node, in order to improve the clustering efficiency and to improve the efficiency of building the cascade tree, in the above step C1, for each level, the similarity between the first labels of the nodes in the level can be determined, and based on the similarity between the first labels, the nodes in the level are clustered to obtain at least one cluster cluster.

[0091] Specifically, nodes with similarity greater than a similarity threshold can be aggregated together to obtain a cluster.

[0092] As another example, for each level, the labels of the nodes in the level are input into the Sbert model to obtain the second vector of the nodes in the level; then, based on the second vector of the nodes in the level and a Gaussian Mixture Model (GMM), the nodes in the level are clustered.

[0093] The Gaussian mixture model is a probability-based clustering model that assumes that data points come from multiple different Gaussian distributions (i.e., clusters), each of which has a different mean and covariance matrix, so that it can better fit different data distributions.

[0094] Specifically, for each level, the main steps of clustering the nodes in the level based on the Gaussian mixture model include: (1) Model initialization. First, set the initial number of Gaussian components (i.e., the number of clusters); then, initialize the parameters of each Gaussian component, including the mean vector, covariance matrix, and mixing coefficient.

[0095] (2) Expectation-maximization algorithm. First, the E-step (Expectation step) is executed: the posterior probability of each node belonging to each Gaussian component is calculated; then, the M-step (Maximization step) is executed: based on the results of the E-step, the parameters of each Gaussian component are re-estimated to maximize the likelihood function.

[0096] (3) Iterative update. Alternate between E-step and M-step until the model parameters converge, that is, the change in the parameters is less than a preset threshold.

[0097] To determine the optimal number of clusters, the present embodiment uses the Bayesian Information Criterion (BIC) to evaluate different numbers of clusters. BIC balances the model's fit and complexity to find the optimal number of clusters for constructing nodes in the next level. Specifically, the main steps for estimating the number of clusters based on BIC include: (1) Log-Likelihood Estimation (LLE): Based on each Gaussian mixture model, calculate its log-likelihood value, that is, the degree of fit of the model to the current data.

[0098] (2) Calculate the number of model parameters: For each high-speed component, calculate the number of its parameters, including the mean vector, covariance matrix, and mixing coefficient.

[0099] (3) Determine the BIC value: ,in, Indicates the number of four fingers, represents the number of parameters of the model, Indicates the number of nodes in this level.

[0100] (4) Select the model with the smallest BIC value: Calculate the corresponding BIC value for different numbers of clusters; then, select the model with the smallest BIC value to obtain the optimal number of clusters.

[0101] The nodes in each level are clustered using the Gaussian mixture model, and the optimal number of clusters is determined using BIC, ultimately forming multiple clusters with semantic relevance. The topics corresponding to the nodes in each cluster are highly similar in semantics and can represent a higher-level topic, which lays the foundation for the subsequent construction of the cascade tree.

[0102] Step C2: for each cluster, create a node corresponding to the cluster as a node in the previous level of the level.

[0103] For example, Figure 5 As shown in the figure, the first level includes 5 nodes. These 5 nodes are clustered to obtain three clusters. The first cluster includes nodes 3 and 5, the second cluster includes nodes 1, 5 and 4, and the third cluster includes nodes 2 and 3. Furthermore, three nodes in the second level are created, namely nodes 6, 7 and 8, among which node 6 corresponds to the first cluster, node 7 corresponds to the second cluster, and node 8 corresponds to the third cluster.

[0104] Step C3: for each cluster, connect the nodes in the cluster with the nodes in the upper level corresponding to the cluster.

[0105] For example, Figure 5 Taking the clusters and nodes shown in FIG5 as an example, node 6 is connected to nodes 3 and 5, node 7 is connected to nodes 1, 5 and 4, and node 8 is connected to nodes 2 and 3.

[0106] Step C4: for each cluster, based on the labels of the nodes in the cluster, determine the labels of the nodes in the upper layer corresponding to the cluster.

[0107] Specifically, for each cluster, the first labels of the nodes in the previous level are determined based on the first labels of the nodes in the cluster; and the second labels of the nodes in the previous level are determined based on the second labels of the nodes in the cluster. In practical applications, both the first and second labels of the nodes in the previous level can be generated using corresponding prompt words and a large-scale language model.

[0108] For example, Figure 5 As shown, node 8 has a first label and a second label. The first label represents the summary of the topic corresponding to node 8, which is generated based on the first label of node 2 and the first label of node 3; the second label represents the key information of the topic corresponding to node 8, which can be obtained by summarizing the key information of the topic corresponding to node 2 and the key information of the topic corresponding to node 4.

[0109] In the above implementation, in a manner from specific to abstract and from micro to macro, each text block with specific and micro content is first taken as a node in the last level, and then from the last level, semantically similar nodes are aggregated together by clustering, and the nodes in the previous level and the labels of the nodes in the previous level are constructed layer by layer until they can no longer be clustered. In this way, the first text is organized and stored in a cascade tree structure. Under this structure, the content of the first text is gradually refined into individual text blocks according to the most abstract theme, providing data support for the subsequent fast and accurate retrieval of valid information in the text.

[0110] The above describes some implementation methods of the above S308. Of course, it should be understood that the above S308 can also be implemented in other ways, and the present embodiment of the application does not limit this.

[0111] The cascade tree construction method provided in the embodiment of the present application adopts an unsupervised clustering method, which can automatically analyze and organize text content without relying on manual annotation, thereby improving the automation and adaptability of the retrieval system. Secondly, the first text is organized and stored in a cascade tree structure. Under this structure, the content of the first text is gradually refined into individual text blocks according to the most abstract topics. This not only enables users to better grasp the information structure and hierarchical relationships, improving the user experience, but also facilitates the rapid and accurate retrieval of effective information in the text.

[0112] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0113] Based on the same inventive concept, the present application also provides an information retrieval device. Figure 6 , is a structural diagram of an information retrieval device 600 provided in an embodiment of the present application. The device 600 includes: a determination module 610 and a query module 620.

[0114] Determination module 610 is used to determine, in response to receiving a first query request, a first node that matches the first query request from the nodes of the first level of the cascade tree; the cascade tree includes multiple nodes, and the multiple nodes are distributed in multiple levels. The nodes in the upper level are obtained by clustering the nodes in the lower level, the nodes in the last level correspond to the text blocks of the first text, and the nodes in the non-last level correspond to the topics of the text blocks.

[0115] The determining module 610 is further configured to determine a second node matching the first query request from nodes in a layer below the layer where the first node is located.

[0116] The query module 620 is configured to determine a query result corresponding to the first query request based on a text block corresponding to the second node if the determined second node is a node in the last level.

[0117] In another embodiment, the query module is also used to determine the second node as the first node if the determined second node is not a node in the last level, and re-execute the step of determining the second node matching the first query request from the next level of the level where the first node is located.

[0118] In another embodiment, each node has a corresponding label, and the label of a node is used to describe the topic or text block corresponding to the node; the first query request carries the second text; The processing module is used for: Performing embedding processing on the second text to obtain a first vector; For each node in the next layer, embed the label of the node to obtain a second vector of the node; Based on the first vector and the second vector of each node in the next layer, a second node matching the first query request is determined from the nodes in the next layer.

[0119] In another embodiment, when determining, from the nodes in the next layer based on the first vector and the second vector of each node in the next layer, the second node that matches the first query request, the processing module performs the following steps: Embedding the label of the first node to obtain a second vector of the first node; For each node of the next level, project the second vector of the node onto the second vector of the first node to obtain a third vector of the node; and determining a difference between the second vector of the first node and the third vector of the node as a fourth vector of the node; From the nodes at the next level, a node whose similarity between the fourth vector and the first vector is greater than or equal to a similarity threshold is determined as a second node matching the first query request.

[0120] In another embodiment, the information retrieval device further comprises: a segmentation module, configured to perform semantic segmentation on the first text to obtain a plurality of text blocks; a creation module, configured to create a plurality of nodes corresponding one-to-one to the plurality of text blocks as nodes in the final layer, and determine labels of the nodes in the final layer based on the text blocks corresponding to the nodes in the final layer; The clustering module is used to cluster the nodes in each level based on the labels of the nodes in each level, starting from the last level, to obtain the nodes in the previous level and the labels of the nodes in the previous level.

[0121] In another embodiment, the segmentation module includes: Combining two adjacent sentences in the first text to obtain a plurality of sentence pairs; For each sentence pair, determining a first score for the sentence pair based on a similarity between two sentences in the sentence pair; For each sentence in the first text, determining a second score for the sentence based on a first score of a first sentence pair in which the sentence is included and a first score of a second sentence pair adjacent to the first sentence pair; the second score represents a degree of semantic difference between the sentence and the adjacent sentence, the first sentence pair comprising the sentence and the sentence following the sentence; Based on a first sentence in the first text having a second score greater than or equal to a score threshold, the first text is divided into a plurality of text blocks.

[0122] In another embodiment, the segmentation module performs the following steps when determining the second score of the sentence based on the first score of the first sentence pair in which the sentence is located and the first score of the second sentence pair adjacent to the first sentence pair: Obtaining a difference between the first score of the second sentence pair and the first score of the first sentence pair to obtain a third score of the sentence; A second score of the sentence is obtained based on an average of the third score of the sentence and third scores of second sentences within a neighborhood of the sentence.

[0123] Obviously, the information retrieval device provided in the embodiment of the present application can be used as Figure 1 The execution subject of the information retrieval method shown is, for example, Figure 1 In the information retrieval method shown in FIG. 1 , steps S102 and S104 can be performed by Figure 6 The information retrieval device shown in FIG. 1 is executed by the determination module 610, and step S106 can be performed by Figure 6 The query module 620 in the information retrieval device shown is executed.

[0124] According to another embodiment of the present application, Figure 6 The various modules in the information retrieval device shown can be individually or all combined into one or several other modules to form a whole, or one (or some) of the modules can be further divided into multiple functionally smaller modules to form a whole, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one module can also be implemented by multiple modules, or the functions of multiple modules can be implemented by one module. In the embodiments of the present application, the information retrieval device may also include other modules. In actual applications, these modules can also be implemented with the assistance of other modules, and can be implemented by the collaboration of multiple modules.

[0125] According to another embodiment of the present application, a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements can be run to execute the following operations: Figure 1 A computer program (including program code) for each step involved in the corresponding method shown in FIG. Figure 6 The information retrieval device shown in the figure and the information retrieval method according to the embodiment of the present application are implemented. The computer program can be recorded on a computer-readable storage medium, for example, and transferred to an electronic device through the computer-readable storage medium and run therein.

[0126] Figure 7 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 7 At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.

[0127] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0128] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.

[0129] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming an information retrieval device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: In response to a first query request, determining a first node that matches the first query request from nodes in a first level of a cascade tree; the cascade tree includes a plurality of nodes, the plurality of nodes being distributed in a plurality of levels, the nodes in a previous level being obtained by clustering nodes in a next lower level, the nodes in the last level corresponding to a text block of the first text, and the nodes in a non-last level corresponding to a topic of the text block; Determining, from a level below the level where the first node is located, a second node that matches the first query request; If the determined second node is a node in the last level, a query result corresponding to the first query request is determined based on the text block corresponding to the second node.

[0130] The above application Figure 1 The methods performed by the information retrieval device disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits within the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0131] The electronic device may also perform Figure 1 Method, and realize information retrieval device in Figure 1 、 Figure 3 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0132] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0133] The embodiment of the present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 1 The method of the embodiment shown is specifically used to perform the following operations: In response to a first query request, determining a first node that matches the first query request from nodes in a first level of a cascade tree; the cascade tree includes a plurality of nodes, the plurality of nodes being distributed in a plurality of levels, the nodes in a previous level being obtained by clustering nodes in a next lower level, the nodes in the last level corresponding to a text block of the first text, and the nodes in a non-last level corresponding to a topic of the text block; Determining, from a level below the level where the first node is located, a second node that matches the first query request; If the determined second node is a node in the last level, a query result corresponding to the first query request is determined based on the text block corresponding to the second node.

[0134] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps in the information retrieval method provided in the embodiment of the present application.

[0135] In short, the above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

[0136] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0137] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0138] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0139] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

Claims

1. An information retrieval method, characterized in that: include: In response to a first query request, determining a first node matching the first query request from nodes in a first level of the cascade tree; The cascade tree includes a plurality of nodes, the plurality of nodes being distributed in a plurality of levels, the nodes in the upper level being obtained by clustering the nodes in the lower level, the nodes in the last level corresponding to the text block of the first text, and the nodes in the non-last level corresponding to the topic of the text block; Determining, from a level below the level where the first node is located, a second node that matches the first query request; If the determined second node is a node in the last level, a query result corresponding to the first query request is determined based on the text block corresponding to the second node.

2. The method according to claim 1, characterized in that The method further comprises: If the determined second node is not a node in the last level, the second node is determined as the first node, and the step of determining a second node matching the first query request from a level below the level where the first node is located is re-executed.

3. The method according to claim 1, characterized in that Each node has a corresponding label, and a label of a node is used to describe the topic or text block corresponding to the node; the first query request carries the second text; The determining, from a level below the level where the first node is located, a second node matching the first query request includes: Performing embedding processing on the second text to obtain a first vector; For each node in the next layer, embed the label of the node to obtain a second vector of the node; Based on the first vector and the second vector of each node in the next layer, a second node matching the first query request is determined from the nodes in the next layer.

4. The method according to claim 3, characterized in that The determining, based on the first vector and the second vector of each node in the next layer, a second node matching the first query request from the nodes in the next layer includes: Embedding the label of the first node to obtain a second vector of the first node; For each node of the next level, project the second vector of the node onto the second vector of the first node to obtain a third vector of the node; and determining a difference between the second vector of the first node and the third vector of the node as a fourth vector of the node; From the nodes at the next level, a node whose similarity between the fourth vector and the first vector is greater than or equal to a similarity threshold is determined as a second node matching the first query request.

5. The method according to claim 1, wherein The method further comprises: Performing semantic segmentation on the first text to obtain multiple text blocks; Creating a plurality of nodes corresponding one-to-one to the plurality of text blocks as nodes in a final level, and determining labels of the nodes in the final level based on the text blocks corresponding to the nodes in the final level; Starting from the last level, the nodes in each level are clustered based on the labels of the nodes in each level to obtain the nodes in the previous level and the labels of the nodes in the previous level.

6. The method according to claim 5, characterized in that The semantic segmentation of the first text is performed to obtain multiple text blocks, including: Combining two adjacent sentences in the first text to obtain a plurality of sentence pairs; For each sentence pair, determining a first score for the sentence pair based on a similarity between two sentences in the sentence pair; For each sentence in the first text, determining a second score for the sentence based on a first score of a first sentence pair in which the sentence is included and a first score of a second sentence pair adjacent to the first sentence pair; the second score represents a degree of semantic difference between the sentence and the adjacent sentence, the first sentence pair comprising the sentence and the sentence following the sentence; Based on a first sentence in the first text having a second score greater than or equal to a score threshold, the first text is divided into a plurality of text blocks.

7. The method according to claim 6, characterized in that Determining the second score of the sentence based on the first score of the first sentence pair in which the sentence is located and the first score of the second sentence pair adjacent to the first sentence pair includes: Obtaining a difference between the first score of the second sentence pair and the first score of the first sentence pair to obtain a third score of the sentence; A second score of the sentence is obtained based on an average of the third score of the sentence and third scores of second sentences within a neighborhood of the sentence.

8. An information retrieval device, characterized in that: include: a determining module configured to, in response to receiving a first query request, determine a first node matching the first query request from nodes at a first level of the cascade tree; The cascade tree includes a plurality of nodes, the plurality of nodes being distributed in a plurality of levels, the nodes in the upper level being obtained by clustering the nodes in the lower level, the nodes in the last level corresponding to the text block of the first text, and the nodes in the non-last level corresponding to the topic of the text block; The determining module is further configured to determine a second node matching the first query request from nodes in a layer below the layer where the first node is located; The query module is configured to determine a query result corresponding to the first query request based on a text block corresponding to the second node if the determined second node is a node in the last level.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the information retrieval method according to any one of claims 1 to 7.

10. A computer program product, characterized in that The computer program product includes a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute part or all of the steps in the information retrieval method according to any one of claims 1 to 7.