Content retrieval method, device and equipment based on machine learning model

By adding identification to the content object and using machine learning models to output indexes, the target content fragments are extracted from the content object, and the problems of delay and cost in traditional search methods are solved, achieving more efficient and accurate information retrieval.

CN120277070APending Publication Date: 2025-07-08BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510387285.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In traditional content retrieval methods, machine learning models have high latency and usage costs in information retrieval, and may not be able to completely recall content related to user queries or recall unrelated content.

Method used

By dividing the content object into multiple content fragments and adding multiple sets of identifiers to each fragment, the machine learning model is used to output the index of the target content fragment, and then extract the target content fragment from the content object to answer the query request.

Benefits of technology

It reduces the number of output word elements of the machine learning model, reduces the search delay and cost, and improves the accuracy of the search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277070A_ABST
    Figure CN120277070A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a content retrieval method, device and equipment based on a machine learning model. The method comprises: in response to a received query request, determining a first content object associated with the query request, the first content object comprising a plurality of content segments and a plurality of groups of identifiers respectively corresponding to the plurality of content segments, and each group of identifiers at least comprising a first index corresponding to the content segment; providing the query request and the first content object to a machine learning model to obtain at least one second index output by the machine learning model, the at least one second index respectively indicating at least one target content segment in the plurality of content segments; and based on the at least one second index and the multiple groups of identifiers, extracting at least one target content fragment from the first content object to respond to the query request.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to a content retrieval method, apparatus, device, computer-readable storage medium, and computer program product based on a machine learning model. Background Art

[0002] Artificial intelligence technology has been widely applied to various information retrieval scenarios. For example, a machine learning model can be used to retrieve required information from content objects such as text and tables. Due to its powerful natural language understanding ability and reasoning ability, the machine learning model can accurately understand the user's retrieval intention and thus provide accurate retrieval results. Summary of the Invention

[0003] In a first aspect of the present disclosure, there is provided a content retrieval method based on a machine learning model. The method includes: in response to receiving a query request, determining a first content object associated with the query request, the first content object including a plurality of content segments and multiple groups of identifiers respectively corresponding to the plurality of content segments, each group of identifiers including at least a first index of the corresponding content segment; providing the query request and the first content object to the machine learning model to obtain at least one second index output by the machine learning model, the at least one second index respectively indicating at least one target content segment among the plurality of content segments; and based on the at least one second index and the multiple groups of identifiers, extracting at least one target content segment from the first content object to answer the query request.

[0004] In a second aspect of the present disclosure, there is provided an apparatus for content retrieval based on a machine learning model. The apparatus includes: a determination module configured to, in response to receiving a query request, determine a first content object associated with the query request, the first content object including a plurality of content segments and multiple groups of identifiers respectively corresponding to the plurality of content segments, each group of identifiers including at least a first index of the corresponding content segment; a provision module configured to provide the query request and the first content object to the machine learning model to obtain at least one second index output by the machine learning model, the at least one second index respectively indicating at least one target content segment among the plurality of content segments; and an extraction module configured to, based on the at least one second index and the multiple groups of identifiers, extract at least one target content segment from the first content object to answer the query request.

[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. Computer-executable instructions are stored on the computer-readable storage medium, and the computer-executable instructions can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided, including computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0008] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram showing an example environment in which the embodiments according to the present disclosure can be implemented;

[0011] Figure 2 A flowchart showing the process of content retrieval based on a machine learning model according to some embodiments of the present disclosure;

[0012] Figure 3 A schematic diagram showing an example of a tree structure according to some embodiments of the present disclosure;

[0013] Figure 4 A schematic structural block diagram showing an example device for content retrieval based on a machine learning model according to some embodiments of the present disclosure; and

[0014] Figure 5 A block diagram showing an electronic device capable of implementing multiple embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0016] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0017] In this article, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.

[0018] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0019] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to the relevant laws and regulations.

[0020] For example, when receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information, so that the user can autonomously choose whether to provide personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0021] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0022] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0023] As used herein, the term "model" can learn the association relationship between corresponding inputs and outputs from training data, so that after training is completed, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably herein.

[0024] A "neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and it generally includes an input layer and an output layer as well as one or more hidden layers between the input layer and the output layer. The neural networks used in deep learning applications usually include many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is used as the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.

[0025] As mentioned above, with the rapid development of artificial intelligence technology, machine learning models can be used to retrieve the required information from content objects such as documents and tables. In traditional technologies, the content in the content object is usually divided into multiple content blocks. A machine learning model is used to perform feature matching based on the user query and the content blocks, so as to recall one or more content blocks that match the user query. The output latency and usage cost of the machine learning model are positively correlated with the number of tokens output. If there are many content blocks associated with the user query, the latency and usage cost will increase significantly. Traditional content block division is usually based on content length or data volume. During the retrieval process, it may occur that not all the content related to the user query is recalled, or the recalled content blocks contain irrelevant content. Thus, it can be seen that the retrieval strategy for using machine learning models to retrieve information in traditional technologies still needs to be improved, and the retrieval cost still needs to be reduced.

[0026] In view of this, embodiments of the present disclosure propose an improved solution for content retrieval based on a machine learning model. In this solution, if a query request is received, a first content object associated with the query request is determined. The first content object includes a plurality of content segments and multiple groups of identifiers respectively corresponding to the multiple content segments, and each group of identifiers includes at least a first index of the corresponding content segment. The query request and the first content object are provided to the machine learning model to obtain at least one second index output by the machine learning model. The at least one second index respectively indicates at least one target content segment among the multiple content segments. Then, based on the at least one second index and the multiple groups of identifiers, at least one target content segment is extracted from the first content object to respond to the query request.

[0027] In embodiments of the present disclosure, a machine learning model is used to determine target content segments that match a query request. However, the machine learning model does not directly output the retrieved target content segments, but outputs the indices of the target content segments. Then, the target content segments are extracted from the content object using the indices to respond to the query request. Thus, the understanding ability and reasoning ability of the machine learning model are fully utilized to ensure the accuracy of the retrieval results. Moreover, the number of tokens that the machine learning model needs to output can be significantly reduced, which is beneficial to reducing the retrieval latency and retrieval cost.

[0028] The following further describes various example implementations of this solution in detail with reference to the accompanying drawings.

[0029] Example environment

[0030] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, an application 120 is installed in the terminal device 110. The user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110.

[0031] In some embodiments of the present disclosure, the application 120 can be any suitable application with an information query function. For example, the application 120 can provide a digital assistant for information query. The digital assistant supports the user 140 to use query requests in text, voice, or other modalities. In some embodiments, if the application 120 is in an active state, the terminal device 110 can present a user interface 150 of the application 120. The user interface 150 can include various types of pages that the application 120 can provide, such as a dialogue page between the user and the digital assistant, a presentation interface of the retrieved content, and so on. In some embodiments, the terminal device 110 can play voice in the user interface 150, and the voice can be, for example, the voice corresponding to the retrieved text content.

[0032] In some embodiments, the application 120 or the digital assistant therein can utilize the machine learning model 160 (which can include one or more machine learning models, for example, can include machine learning models 160-1, machine learning models 160-2, ……, machine learning models 160-N, etc., where N is a positive integer. For convenience of description, one or more machine learning models are collectively referred to as the machine learning model 160 in this article) to support the interaction with the user 140. For example, the application 120 or the digital assistant therein can utilize one or more machine learning models 160 to provide an information query service to the user 140. In some embodiments, the terminal device 110 can communicate with the server 130 to implement the supply of the services of the application 120. As Figure 1 shown, the server 130 can call the machine learning model 160 to support the application 120 in providing an information query service to the user 140 based on the output of the machine learning model 160.

[0033] In some embodiments, the machine learning model 160 can be different types of models. In some embodiments, one or more machine learning models 160 can be constructed based on a language model (LM). The machine learning model used is a content generation model that can generate a corresponding output based on the model input. In some embodiments, the machine learning model based on the language model can process model inputs in text modality (e.g., natural language and / or machine language) and / or non-text modality (e.g., images, speech, videos, etc.), and can generate a desired output according to the model input and the prompt. Here, the prompt is used to guide the machine learning model to generate an answer that can solve the user query indicated by the model input. In the application scenario for supporting user conversations, the input of the user 140 can be provided to the machine learning model 160 as at least a part of the model input (the other part can include the prompt). This user input is regarded as a question or a query request. Based on the model output, a corresponding response can be provided to the user 140.

[0034] In some embodiments, one or more machine learning models 160 can be voice-related models, including an automatic speech recognition (ASR) model and a text-to-speech (TTS) model. The input of the ASR model is speech, and the output is text. The input of the TTS model is text, and the output is the corresponding speech.

[0035] In some embodiments, the terminal device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 is also capable of supporting any type of user interface (such as a "wearable" circuit, etc.).

[0036] In some embodiments, the server 130 may include, but is not limited to, mainframes, edge computing nodes, computing devices in a cloud environment, etc. It may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The server may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and so on.

[0037] It should be understood that the structures and functions of the various elements in the environment 100 are described only for exemplary purposes, without implying any limitation on the scope of the present disclosure.

[0038] Example process

[0039] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings. Figure 2 FIG. 200 is a flowchart showing a process 200 of content retrieval based on a machine learning model according to some embodiments of the present disclosure. Hereinafter, for ease of discussion, the execution of the process 200 is described from the perspective of the server 130, but this is only exemplary. The process 200 may also be executed by the terminal device 110 or jointly executed by the terminal device 110 and the server 130.

[0040] In block 210, if a query request is received, the server 130 determines a first content object associated with the query request. The first content object includes a plurality of content segments and multiple groups of identifiers, and the multiple groups of identifiers respectively indicate the plurality of content segments. In some embodiments, the first content object may include text, tables, images, audio, video, or a combination of one or more of them, and so on. For example, the first content object may be a document, and the document may include text, tables, images, and so on.

[0041] Each group of identifiers includes at least a first index indicating the corresponding content segment. In the case where the modality of the first content object is different, the content segment and the first index may also be different. As an example, the first content object may include text, and the multiple content segments may include multiple text segments. The first index may include various data contents indicating the corresponding text segments, such as words, symbols, labels, and so on. As another example, the first content object may also include audio, the content segments may include audio segments, and the first index may include, for example, timestamps. Of course, the types of the above first content object and the first index are both exemplary, and any other appropriate first content object and any appropriate first index can be selected, and the embodiments of the present disclosure do not limit this.

[0042] In some embodiments, the server 130 may, in response to receiving a query request from the terminal device 110, query the associated content object from a predetermined data source (such as a database). For example, the server 130 may query the content object associated with the query request of the user 140 from the database associated with the application 120. Of course, the server 130 may also obtain the content object from other data sources. For example, the server 130 may also obtain the content object provided by the user 140.

[0043] In some embodiments, the server 130 may, based on the query request, obtain a content object in a predetermined format from a predetermined data source. The content object in the predetermined format may include at least some of the identifiers in the multiple groups of identifiers. For example, the content object in the predetermined format may include the first index in each group of identifiers. In some embodiments, the first content object may include a document in a predetermined format. The document in the predetermined format may be, for example, a document in a markup language, such as a Markdown document. As an example, the database associated with the application 120 may store a Markdown document, and the Markdown document may include at least some of the identifiers indicating the text segments.

[0044] In some embodiments, the content object obtained by server 130 may not include an identifier for the content segment either. Server 130 may add an identifier to the obtained content object to obtain a first content object. Specifically, server 130 may determine the original content object associated with the query request in response to the query request. Server 130 may identify multiple content segments from the original content object (for example, multiple content segments may be identified based on the relevance between contents or the content structure). After that, server 130 may add multiple groups of identifiers to the original content object based on the identification result to obtain a first content object. In this way, even if the original content object does not have an identifier added, by adding an identifier to the original content object, the required content can be retrieved from it, which is beneficial to improving flexibility and can expand the application scenarios. The modality of the original content object here may be the same as or different from that of the first content object. For example, both the original content object and the first content object may include text. Also, for example, the original content object may be a table or audio, and server 130 may convert the table or audio into text.

[0045] The following provides an exemplary description of the specific content included in the identifier and the process of adding the identifier in combination with some specific examples, but it should be understood that this is only exemplary. In actual application scenarios, the identifier is not limited to including the content shown below, nor is it limited to being added to the content object in the manner shown below.

[0046] In some embodiments, multiple content segments may be configured to include multi-level content segments, and the first index may include a level identifier indicating the level of the corresponding content segment. Server 130 may determine the level of the first content segment among the multiple content segments. After that, server 130 may add a level identifier corresponding to the level of the first content segment to the original content object as at least a part of the first index of the first content segment. The first content segment may be any one of the content segments in the original content object.

[0047] As an example, the original content object and the first content object may be documents. The original content object may include the original document shown in the following table.

[0048] Table 1

[0049]

[0050] Server 130 can identify multiple text segments in the original document in Table 1. Server 130 can determine the levels of the multiple text segments based on predetermined rules (such as rules based on document structure, or rules based on the correlation between contents, etc.). For example, server 130 can determine the text content to which "Title 1" belongs as a first-level text segment, determine the text content to which "Title 1.1", "Title 1.2", and "Title 1.3" belong as second-level text segments, determine the text content to which "Title 1.2.1" belongs as a third-level text segment, and so on. After that, server 130 can add level identifiers corresponding to the levels of each text segment to the original document shown in Table 1. For example, characters such as "#", "##", "", etc. are added for the first-level text segment, second-level text segment, and third-level text segment respectively to indicate the levels of the corresponding text segments. Thus, content segments can be efficiently and flexibly divided according to the correlation between contents, and then content matching the query request can be efficiently and flexibly output.

[0051] The document after adding the level identifiers is as shown in the following table.

[0052] Table 2

[0053]

[0054] In some embodiments, content segments can not only have corresponding levels, but also have corresponding orders in a group of content segments with the same level. Server 130 can determine the level of the second content segment among the multiple content segments, and the order of the second content segment in a group of content segments with the same level. After that, server 130 can add an order identifier corresponding to the order of the second content segment to the original content object as at least a part of the first index of the second content segment. Thus, using the first index can accurately identify the order of each content segment, which is beneficial to accurately outputting the retrieval results subsequently.

[0055] As an example, still referring to the documents shown in Table 1 and Table 2, server 130 can determine the content corresponding to "Title 1.2.1.1" and "Title 1.2.1.2" in Table 1 as ordered list segments, and the orders of the ordered list segments corresponding to "Title 1.2.1.1" and "Title 1.2.1.2" are "1" and "2" respectively. Server 130 can also determine the content corresponding to "Title X" in Table 1 as an unordered list segment. After that, server 130 can add order identifiers "1." and "2." to the document shown in Table 1 to indicate the orders of the ordered list segments corresponding to "1. Title 1.2.1.1" and "2. Title 1.2.1.2" respectively. Server 130 can also add an order identifier "-" indicating that the content corresponding to "- Title X" is an unordered list segment to the document shown in Table 1.

[0056] It should also be noted that the level relationships between text segments can be predefined. For example, it can be predetermined that the levels of ordered text segments and unordered text segments are the same, and the levels of ordered text segments and unordered text segments are lower than those of first-level text segments, second-level text segments, and third-level text segments. In this case, the sequence identifiers "1.", "2." can not only indicate that the corresponding list segments belong to ordered list segments and the order of the corresponding list segments, but also can, to a certain extent, indicate the levels of the corresponding list segments. Similarly, "-" can not only indicate that the corresponding list segment belongs to an unordered list segment and that the corresponding list segment has no order, but also can, to a certain extent, indicate the level of the corresponding list segment. It can be seen that in some cases, the sequence identifier can also indicate the level of the corresponding content segment and can, to a certain extent, serve as a level identifier.

[0057] In some embodiments, each group of identifiers includes a start identifier indicating the start position of the corresponding content segment. The server 130 can determine the start position of the third content segment among the multiple content segments. Then, in the original content object, a start identifier is added to the start position of the third content segment to be at least a part of the first index of the third content segment. As an example, the server 130 can determine the start positions of text segments such as "Title 1", "Title 1.1", "Title 1.2" in the document shown in Table 1. Then, the server 130 can add, for example, "#", "##", "" to the start positions of the text segments corresponding to "Identifier 1", "Title 1.1", "Title 1.2" in the document shown in Table 1 to indicate the start positions of these text segments. Of course, the start identifier is not limited to the above identifiers, and the server 130 can also add, for example, "begin" or other characters to the start position of the text segment to indicate the start position of the corresponding text segment.

[0058] In some embodiments, each group of identifiers can also include an end identifier indicating the end position of the corresponding content segment. The server 130 can identify the end positions of the multiple content segments from the original content object. Then, the server 130 can add end identifiers to the end positions corresponding to each of the multiple content segments in the original content object. For example, the server 130 can add characters such as "end" to the end positions of the content segments in the original content object to represent the end positions of the content segments. By adding the end identifier, the end positions of each content segment can be accurately identified, which is conducive to accurately extracting the content segments.

[0059] In some examples, if multiple content segments include multi-level content segments, the server 130 can determine that two adjacent content segments in the original content object are both at the first level, and add a first termination identifier corresponding to the first level to the recognized end position of the previous content segment among the two adjacent content segments. The first level can be any level. Correspondingly, the fact that two adjacent content segments are both at the first level can be understood as that the levels of the two adjacent content segments are the same. That is, if the levels of two adjacent content segments are the same, the server 130 adds a termination identifier corresponding to the level of the previous content segment at the end position of the previous content segment.

[0060] As an example, taking the document shown in Table 2 as an example, the text segment corresponding to "## Title 1.1" and the text segment corresponding to "## Title 1.2" are both second-level text segments. Therefore, the server 130 can add the character "「##end」" as a termination identifier at the end position of the text segment corresponding to "## Title 1.1". Similarly, the text segment corresponding to "1. Title 1.2.1.1" and the text segment corresponding to "2. Title 1.2.1.2" are both ordered list segments. The server 130 can add the character "「1.end」" as a termination identifier at the end position of the ordered list segment corresponding to "1. Title 1.2.1.1". Specifically, as shown in Table 3.

[0061] Table 3

[0062]

[0063] In some examples, if it is determined that the previous content segment among two adjacent content segments in the original content object is at the second level and the subsequent content segment is at the third level higher than the second level, the server 130 can add multiple termination identifiers to the recognized end position of the previous content segment. The multiple termination identifiers at least include a second termination identifier corresponding to the second level and a third termination identifier corresponding to the third level. As an example, continuing to refer to Table 3, since the text segment corresponding to "- Title X" is an unordered list segment and the text segment corresponding to "## Title 1.3" is a second-level text segment, the end position of the text segment corresponding to "- Title X" is also the end position of the text segment corresponding to " Title 1.2.1" and the text segment corresponding to "## Title 1.2". The server 130 can add three termination identifiers, namely "「-end」", "「end」", and "「##end」", at this position.

[0064] In some embodiments, server 130 may convert an original content object into a structured content object that conforms to a predetermined format. The structured content object of the predetermined format may include at least some of the identifiers in multiple groups of identifiers. If the structured content object of the predetermined format contains all of the identifiers in multiple groups of identifiers, server 130 may use the structured content object of the predetermined format as the first content object. If the structured content object of the predetermined format includes some of the identifiers in multiple groups of identifiers, server 130 may also add the remaining identifiers in multiple groups of identifiers to the structured content object to form the first content object. For example, the structured object of the predetermined format may include a first index of each identifier in multiple groups of identifiers, and server 130 may also add multiple termination identifiers corresponding to multiple content segments to the structured object to obtain the first content object.

[0065] As an example, Figure 3 FIG. 300 is a schematic diagram showing an example of a tree structure according to some embodiments of the present disclosure. If server 130 obtains the original document shown in Table 1 according to a query request. Server 130 may identify the levels of multiple text segments in the original document to obtain a tree structure as shown in Figure 3 FIG. 300. Each node in the tree structure may indicate the level of the corresponding text segment, and the tree structure may also indicate the hierarchical relationship between multiple text segments. For example, node 302 may indicate that the text segment corresponding to "Title 1" is a first-level text segment, node 314 indicates that the text segment corresponding to "Title 1.2" is a second-level text segment, node 334 indicates that the content corresponding to "Title 1.2.1.2" is an ordered list segment, and so on.

[0066] Server 130 may add a first index (which may also be referred to as an index name) to each node in the tree structure to obtain the structured document shown in Table 2. The first index may include at least one of, for example, a level identifier, an order identifier, a title of a text segment (or a part of the content at the beginning of the text segment), and the first index may also include the title of the corresponding content segment. For example, the first index corresponding to node 302 includes the level identifier "#" and the title "Title 1" of the corresponding text segment. The first index corresponding to node 334 includes the order identifier "2." and the title "Title 1.2.1.2". It can be seen that server 130 can achieve the purpose of adding at least some identifiers to the original document by converting the original document shown in Table 1 into a structured document such as the Markdown document shown in Table 2.

[0067] Further, the server 130 can also add termination identifiers for each text segment to the structured document shown in Table 2, such as "「-end」", "「end」", "「##end」", etc., to obtain the target document shown in Table 3. Regarding the way of adding termination identifiers, it will not be elaborated here, and reference can be made to the explanations in the foregoing content.

[0068] As another example, if the original document is a structured document in a predetermined format such as a Markdown document, the original document actually already contains some of the identifiers in multiple groups of identifiers. In this case, the server 130 can add the remaining identifiers in the multiple groups of identifiers to the original document to form the first content object. For example, if the Markdown document shown in Table 2 is stored in the database related to the application 120, the server 130 only needs to add the termination identifiers as shown above to the Markdown document to obtain the first content object shown in Table 3.

[0069] In some embodiments, the server 130 can generate prompt word information for the machine learning model 160-1 based on the original content object. The server 130 can provide the prompt word information to the machine learning model 160-1 to obtain the model output generated by the machine learning model 160-1. After that, the server 130 can determine the first content object based on the model output. That is, the server 130 can also add identifiers indicating content segments to the original content object with the help of the machine learning model 160-1 to improve the efficiency and accuracy of adding identifiers. It should be noted that the above identifiers and the way of adding identifiers are only exemplary, and any appropriate identifiers can be selected according to actual needs, and the identifiers can be added to the original content object in any appropriate way. The embodiments of the present disclosure do not limit this.

[0070] Returning to process 200, at block 220, the server 130 can provide the query request and the first content object to the machine learning model 160-2 to obtain at least one second index output by the machine learning model 160-2. The second index respectively indicates at least one target content segment among multiple content segments. In some embodiments, the server 130 can generate prompt word information for the machine learning model based on the query request, the first content object, and the description of multiple groups of identifiers. The server 130 can provide the prompt word information to the machine learning model 160-1 to obtain the model output generated by the machine learning model 160-2. After that, the server 130 can determine at least one second index based on the model output. The at least one second index can respectively indicate at least one target content segment matching the query request.

[0071] The description of multiple groups of identifiers can be understood as an explanation and illustration of the identifiers added to the content segments, so as to facilitate the machine learning model 160-2 to understand the division method of multiple content segments in the first content object and the relationship between multiple content segments (such as hierarchical relationship, sequential relationship, etc.). As an example, for the target document shown in Table 3, the description of multiple groups of identifiers can include, for example, "the level identifier '#' for the first-level text segment, the level identifier '##' for the second-level text segment, the level identifier '' for the third-level text segment, the sequential identifier '-' for the unordered list segment, the sequential identifiers '1.', '2.' for the ordered list segment, and the corresponding termination identifiers '#end', '##end', 'end', '1.end', '2.end', '-end', etc. to distinguish text segments". It should be understood that the above description of multiple groups of identifiers is only exemplary. In actual application scenarios, multiple groups of identifiers can be described by any appropriate statements, and the embodiments of the present disclosure do not limit the description method of identifiers.

[0072] In some embodiments, the second index may include all or part of the content of the first index. As an example, the machine learning model 160-2 may be configured to output the same content as the first index of the target content segment as the second index. In some examples, the second index may include at least one of the level identifier, sequential identifier, and start identifier of the corresponding target content segment, and the second index may also include the title of the corresponding target content segment. As an example, as Figure 3 shown, the machine learning model 160-2 may be configured to output the index names of each node in the tree structure, such as "##Title 1.2", "1.Title 1.2.1.1", "-Title X", etc.

[0073] In some embodiments, if multiple content segments in the first content object include multi-level content segments, and at least one of the target content segments includes a parent content segment and a child content segment belonging to the parent content segment, the machine learning model 160-2 may be configured to output the second index corresponding to the parent content segment, rather than outputting the second index of the child content segment. Thus, content duplication can be avoided.

[0074] Combined with Figure 2As shown, at block 230, server 130 extracts at least one target content segment from the first content object based on at least one second index and multiple sets of identifiers to respond to a query request. In some embodiments, server 130 may determine at least one set of identifiers in the multiple sets of identifiers that match the at least one second index. Thereafter, server 130 may extract at least one target content segment from the first content object based on the at least one set of identifiers. As an example, server 130 may use a regular expression based on the at least one second index to determine at least one set of identifiers that match the at least one second index. Of course, the identifiers that match the second index may also be determined by any other suitable means, and the embodiments of the present disclosure are not limited thereto.

[0075] In some embodiments, server 130 may determine at least one first index that matches the at least one second index and at least one termination identifier corresponding to the at least one first index from the multiple sets of identifiers of the first content object based on the at least one second index. Thereafter, based on the at least one first index and the at least one termination identifier, the at least one target content segment is extracted from the first content. As an example, as shown in Table 3, if machine learning model 160-2 outputs the second index "##Title 1.1", server 130 may determine the first index "##Title 1.1" in the document shown in Table 3 and the termination identifier "##end" corresponding to the first index. Thereafter, server 130 may extract the corresponding secondary text segment from the document based on the first index "##Title 1.1" and the termination identifier "##end".

[0076] In some examples, server 130 may feedback the at least one target content segment to terminal device 110, and terminal device 110 may present the at least one target content segment through user interface 150 of application 120 to respond to the query request of user 140. In other examples, server 130 may generate speech corresponding to the at least one target content segment, and server 130 may provide the speech to terminal device 110, and terminal device 110 may also play the speech using application 120 to respond to the query request of user 140. Of course, server 130 may also provide the retrieval result to user 140 by other means.

[0077] In this way, in the embodiments of the present disclosure, a machine learning model is used to determine a target content segment that matches a query request. However, the machine learning model does not directly output the retrieved target content segment, but outputs an index of the target content segment. Then, the index is used to extract the target content segment from the content object to answer the query request. Thus, the understanding ability and reasoning ability of the machine learning model are fully utilized to ensure the accuracy of the retrieval result. Moreover, the number of tokens that the machine learning model needs to output can be significantly reduced, which is beneficial to reducing the retrieval latency and retrieval cost.

[0078] Example device and equipment

[0079] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. Figure 4 FIG. shows a schematic structural block diagram of an exemplary apparatus 400 for content retrieval based on a machine learning model according to some embodiments of the present disclosure. The apparatus 400 can be implemented as or included in a server 130. Each module / component in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.

[0080] As Figure 4 shown, the apparatus 400 includes: a determination module 410 configured to, in response to receiving a query request, determine a first content object associated with the query request, the first content object including a plurality of content segments and multiple groups of identifiers respectively corresponding to the plurality of content segments, each group of identifiers including at least a first index of the corresponding content segment; a providing module 420 configured to provide the query request and the first content object to a machine learning model to obtain at least one second index output by the machine learning model, the at least one second index respectively indicating at least one target content segment among the plurality of content segments; and an extraction module 430 configured to, based on the at least one second index and the multiple groups of identifiers, extract at least one target content segment from the first content object to answer the query request.

[0081] In some embodiments, the determination module 410 is further configured to: identify a plurality of content segments from an original content object associated with the query request; and based on the identification result, add multiple groups of identifiers to the original content object to obtain the first content object.

[0082] In some embodiments, the plurality of content segments include multi-level content segments, and the determination module 410 is further configured to: determine the level of a first content segment among the plurality of content segments; and in the original content object, add a level identifier corresponding to the level of the first content segment as at least a part of the first index of the first content segment.

[0083] In some embodiments, the determining module 410 is further configured to: determine the order of a second content segment among a plurality of content segments in a group of content segments at the same level; and in the original content object, add an order identifier corresponding to the order of the second content segment as at least a part of the first index of the second content segment.

[0084] In some embodiments, the determining module 410 is further configured to: determine the starting position of a third content segment among a plurality of content segments; and in the original content object, add a corresponding starting identifier to the starting position of the third content segment as at least a part of the first index of the third content segment.

[0085] In some embodiments, each group of identifiers further includes a termination identifier, and the determining module 410 is further configured to: in the original content object, add termination identifiers to the respectively identified ending positions of the plurality of content segments.

[0086] In some embodiments, the plurality of content segments include multi-level content segments, and the determining module 410 is further configured to perform at least one of the following: in response to determining that two adjacent content segments in the original content object are both at the first level, add a first termination identifier corresponding to the first level to the identified ending position of the previous content segment of the two adjacent content segments, or in response to determining that the previous content segment of two adjacent content segments in the original content object is at the second level and the subsequent content segment is at the third level higher than the second level, add a plurality of termination identifiers to the identified ending position of the previous content segment, the plurality of termination identifiers at least including a second termination identifier corresponding to the second level and a third termination identifier corresponding to the third level.

[0087] In some embodiments, the providing module 420 is further configured to: generate prompt word information for a machine learning model based on a query request, a first content object, and a description of multiple groups of identifiers; and provide the prompt word information to the machine learning model to obtain a model output indicating at least one second index from the machine learning model.

[0088] In some embodiments, the extracting module 430 is further configured to: determine at least one group of identifiers that match at least one second index among multiple groups of identifiers; and based on the at least one group of identifiers, extract at least one target content segment from the first content object.

[0089] In some embodiments, each group of identifiers includes a first index of a corresponding content segment and a termination identifier corresponding to the first index, the termination identifier indicating the end position of the corresponding content segment, and the extraction module 430 is further configured to extract at least one target content segment from the first content object based on at least one first index that matches at least one second index and at least one termination identifier corresponding to the at least one first index.

[0090] The units and / or modules included in the apparatus 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units and / or modules in the apparatus 400 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0091] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 5 The illustrated electronic device 500 is merely exemplary and should not impose any limitation on the functionality and scope of the embodiments described herein. Figure 5 The illustrated electronic device 500 can include or be implemented as Figure 1 the server 130 in Figure 4 the apparatus 400.

[0092] As Figure 5 shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 can include, but are not limited to, one or more processors 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processor 510 can be an actual or virtual processor and is capable of performing various processes according to the executable instructions stored in the memory 520. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.

[0093] An electronic device 500 generally includes multiple computer storage media. Such media can be any accessible media that the electronic device 500 can access, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a magnetic disk, or any other medium that can be capable of storing information and / or data and can be accessed within the electronic device 500.

[0094] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 that has one or more executable instruction modules that are configured to execute the various methods or actions of the various embodiments of the present disclosure.

[0095] The communication unit 540 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented by a single computing cluster or multiple computer machines that can communicate through a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.

[0096] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) as needed through the communication unit 540, external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device that enables the electronic device 500 to communicate with one or more other electronic devices (such as a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0097] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer-executable instruction product is also provided. The computer-executable instruction product is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0098] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer-executable instruction products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable executable instructions.

[0099] These computer-executable instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-executable instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0100] The computer-executable instructions can be loaded onto a computer, other programmable data processing device, or other device, such that a series of operation steps are executed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-executable instruction products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, an executable instruction, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0102] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of technologies in the market, or to enable other ordinary skill in the art to understand the various implementation manners disclosed herein.

Claims

1. A content retrieval method based on a machine learning model, comprising: In response to receiving a query request, determining a first content object associated with the query request, the first content object including a plurality of content segments and multiple groups of identifiers respectively corresponding to the plurality of content segments, each group of identifiers including at least a first index of the corresponding content segment; Providing the query request and the first content object to a machine learning model to obtain at least one second index output by the machine learning model, the at least one second index respectively indicating at least one target content segment among the plurality of content segments; And Based on the at least one second index and the multiple groups of identifiers, extracting the at least one target content segment from the first content object to answer the query request.

2. The method according to claim 1, wherein determining the first content object includes: Identifying the plurality of content segments from an original content object associated with the query request; And Based on the result of the identification, adding the multiple groups of identifiers to the original content object to obtain the first content object.

3. The method according to claim 2, wherein the plurality of content segments include multi-level content segments, and adding the multiple groups of identifiers to the original content object includes: Determining the level of a first content segment among the plurality of content segments; And In the original content object, adding a level identifier corresponding to the level of the first content segment as at least a part of the first index of the first content segment.

4. The method according to claim 2, wherein adding the multiple groups of identifiers to the original content object includes: Determining the order of a second content segment among the plurality of content segments in a group of content segments with the same level; And In the original content object, adding an order identifier corresponding to the order of the second content segment as at least a part of the first index of the second content segment.

5. The method according to claim 2, wherein adding the multiple groups of identifiers to the original content object includes: Determining the starting position of a third content segment among the plurality of content segments; And In the original content object, adding a corresponding starting identifier to the starting position of the third content segment as at least a part of the first index of the third content segment.

6. The method according to claim 2, wherein each group of identifiers further includes a termination identifier, and adding the multiple groups of identifiers to the original content object includes: In the original content object, adding the termination identifier to the respectively identified corresponding end positions of the plurality of content segments.

7. The method according to claim 6, wherein the plurality of content segments include multi-level content segments, and adding the termination identifier to the respectively identified corresponding end positions of the plurality of content segments includes at least one of the following: In response to determining that two adjacent content segments in the original content object are both at the first level, adding a first termination identifier corresponding to the first level to the identified end position of the previous content segment among the two adjacent content segments, or In response to determining that a previous content segment among two adjacent content segments in the original content object is at a second level and a subsequent content segment is at a third level higher than the second level, adding a plurality of termination identifiers to the identified end position of the previous content segment, where the plurality of termination identifiers at least include a second termination identifier corresponding to the second level and a third termination identifier corresponding to the third level.

8. The method according to claim 1, wherein providing the query request and the first content object to a machine learning model includes: Generating prompt word information for the machine learning model based on the query request, the first content object, and the description of the multiple groups of identifiers; And Providing the prompt word information to the machine learning model to obtain a model output indicating the at least one second index from the machine learning model.

9. The method according to claim 1, wherein extracting the at least one target content segment from the first content object includes: Determining at least one group of identifiers among the multiple groups of identifiers that matches the at least one second index; And Extracting the at least one target content segment from the first content object based on the at least one group of identifiers.

10. The method according to claim 9, wherein each group of identifiers includes a first index of a corresponding content segment and a termination identifier corresponding to the first index, the termination identifier indicating the end position of the corresponding content segment, and extracting the at least one target content segment from the first content object based on the at least one group of identifiers includes: Extracting the at least one target content segment from the first content object based on at least one first index that matches the at least one second index and at least one termination identifier corresponding to the at least one first index.

11. An apparatus for content retrieval based on a machine learning model, comprising: A determination module configured to, in response to receiving a query request, determine a first content object associated with the query request, the first content object including a plurality of content segments and multiple groups of identifiers respectively corresponding to the plurality of content segments, and each group of identifiers at least including a first index of a corresponding content segment; A provision module configured to provide the query request and the first content object to a machine learning model to obtain at least one second index output by the machine learning model, where the at least one second index respectively indicates at least one target content segment among the plurality of content segments; And An extraction module configured to extract the at least one target content segment from the first content object based on the at least one second index and the multiple groups of identifiers to answer the query request.

12. An electronic device, comprising: At least one processor; And At least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, and the instructions, when executed by the at least one processor, cause the electronic device to execute the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having computer-executable instructions stored thereon, the computer-executable instructions being executable by a processor to implement the method according to any one of claims 1 to 10.

14. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Neural network calculation method and device, electronic equipment and storage medium

    CN121070445A