Information processing method, device and equipment
By filtering and deduplicating historical dialogues, the problem of inaccurate results from large language models is solved, thus improving the accuracy and efficiency of information processing.
Patent Information
- Application Number
- CN202511565516.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-01-30
AI Technical Summary
When processing user input, large language models may produce inaccurate results due to the presence of incomplete intent recognition and invalid information in historical dialogue information.
By filtering historical dialogues that have not completed intent recognition from historical dialogue information, removing preset characters, and performing deduplication, the target historical dialogue is obtained and input into a large language model for processing.
It improves the accuracy and efficiency of large language models in processing queries, reduces interference from irrelevant information, and enhances the model's data processing capabilities.
Smart Images

Figure CN121434352A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to an information processing method. This specification also relates to an information processing apparatus and a computing device. Background Technology
[0002] With the continuous development of natural language interaction technology, more and more fields can achieve interaction through large language models, reducing the need for excessive manual operation and improving problem-solving efficiency. When processing user input, large language models typically need to identify the user's intent and process the input accordingly. While processing user input, large language models can refer to historical dialogue information; however, this historical information may contain data that is not highly relevant to the user's intent, affecting the processing of user input and leading to inaccurate results.
[0003] Therefore, how to provide a method for accurately processing user input information using large language models is an urgent technical problem to be solved. Summary of the Invention
[0004] In view of this, one or more embodiments of this specification provide an information processing method, apparatus, and device to solve the problem of inaccurate processing results in existing information processing methods.
[0005] According to a first aspect of one or more embodiments of this specification, an information processing method is provided, comprising: The system retrieves the query input by the user on the target interface; the target interface has corresponding historical dialogue information, including the user's historical queries and the historical answers from the large language model for the historical queries. Historical dialogues with incomplete intent identification are filtered from the historical dialogue information; Remove preset characters from the historical dialogues in which intent recognition was not completed to obtain preprocessed historical dialogues; The preprocessed historical dialogues are deduplicated to obtain the target historical dialogues; The target historical dialogue and the query are input into the large language model to obtain the output information of the large language model.
[0006] According to a second aspect of one or more embodiments of this specification, an information processing apparatus is provided, comprising: The query acquisition module is used to acquire the query entered by the user on the target interface; the target interface has corresponding historical dialogue information, which includes the user's historical queries and the historical answers of the large language model to the historical queries; A filtering module is used to filter historical dialogues from the historical dialogue information to obtain those with incomplete intent recognition. The first preprocessing module is used to remove preset characters from the historical dialogues in which intent recognition was not completed, and obtain preprocessed historical dialogues. The second preprocessing module is used to perform deduplication on the preprocessed historical dialogue to obtain the target historical dialogue. The model processing module is used to input the target historical dialogue and the query into the large language model and obtain the output result information of the large language model.
[0007] According to a third aspect of one or more embodiments of this specification, a computing device is provided, including a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein the processor, when executing the computer instructions, implements the steps of the information processing method.
[0008] According to a fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided that stores computer instructions which, when executed by a processor, implement the steps of the information processing method.
[0009] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the information processing method described above.
[0010] One embodiment of this specification can achieve at least the following beneficial effects: By acquiring the query input by the user on the target interface, which may contain corresponding historical dialogue information, it is possible to filter out historical dialogues that have not completed intent recognition from the historical dialogue information, remove preset characters from the historical dialogues that have not completed intent recognition, obtain preprocessed historical dialogues, perform deduplication on the preprocessed historical dialogues, obtain the target historical dialogue, and input the target historical dialogue and the query into a large language model to obtain the output result information of the large language model. On the one hand, it is possible to filter out historical dialogues that have not completed intent recognition from the historical dialogue information, which can avoid inputting historical dialogues that have completed intent recognition into the large language model, thereby improving the accuracy of the large language model in processing queries; on the other hand, it is also possible to remove preset characters from historical dialogues that have not completed intent recognition, thereby reducing information irrelevant to the intent recognition of the large language model, while also performing deduplication, reducing the amount of data input into the large language model, and improving the efficiency of the large language model in processing information. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram illustrating information processing in the prior art, provided by an embodiment of this specification; Figure 2 This is a schematic diagram of the overall architecture of an information processing method provided in one embodiment of this specification; Figure 3 This is a flowchart illustrating an information processing method provided in one embodiment of this specification; Figure 4 This is a schematic diagram of the overall flow of an information processing method provided in one embodiment of this specification; Figure 5 This specification provides an embodiment corresponding to... Figure 2 A schematic diagram of the structure of an information processing device; Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0013] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0014] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.
[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “an,” “an,” “the,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification includes any or all possible combinations of one or more associated listed items.
[0016] The terms “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in the process, method, product, or apparatus that includes said elements is not excluded.
[0017] Although the terms "first," "second," etc., may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, "first" may also be referred to as "second," and similarly, "second" may also be referred to as "first," without departing from the scope of one or more embodiments of this specification. Ordinal numbers such as "first," "second," etc., do not necessarily indicate order; often they are used to facilitate the distinction of objects. For example, "first server" and "second server" usually refer to two servers. To distinguish these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.
[0018] Depending on the context, the word "if" as used here can be interpreted as "when," "when," or "in response to determination."
[0019] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct reception and transmission; it can also mean indirect reception and transmission. For example, when device A receives data sent by device B, it can be understood as device A directly receiving data sent by device B, or it can be understood as device A indirectly receiving data sent by device B through other devices such as device C. Similarly, when device B sends data to device A, it can be understood as device B sending data directly to device A, or it can be understood as device B indirectly sending data to device A through other devices such as device C. Here, device C can be a single device, or it can be a cluster of two or more devices.
[0020] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "structure A is connected to structure B," unless it is explicitly stated that structure A and structure B are directly connected, it should be understood that structure A can be directly connected to structure B, or indirectly connected to structure B. Similarly, when describing "structure A is above structure B," unless it is explicitly stated that structure A is directly above structure B (structure A is adjacent to structure B and structure A is above structure B), it should be understood that structure A can be directly above structure B, or indirectly above structure B (structure A and structure B are separated by other elements, and structure A is above structure B). And so on.
[0021] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation entry points shall be provided for users to choose to authorize or refuse.
[0022] The following explains the terms and concepts used in one or more embodiments of this specification.
[0023] Multi-turn dialogue: In order to achieve a goal, the user engages in a dialogue with the large language model through multiple questions and answers.
[0024] Intent recognition: This is used to determine the true purpose of the user's natural language input and map that purpose to a business category that the system can process. For example, if a user inputs, "I want to take the bus to XXX," the user's intent should be recognized as: querying nearby bus stops.
[0025] To facilitate understanding of the relevant technologies, Figure 1 This is a schematic diagram illustrating information processing in a prior art as provided in this specification. Figure 1As shown, after obtaining the user's input query, a static switch needs to determine whether to retrieve historical dialogues. If so, all historical dialogues can be retrieved, and the user's input query and historical dialogues are input into the large language model so that the large language model can process based on the query and historical dialogues to obtain the processing result. If not, only the user's input query can be input into the large language model so that the large language model can process based on the query to obtain the processing result. However, whether to retrieve historical dialogues is preset and cannot flexibly identify the scenario. This results in historical dialogues either being effective for all inputs, meaning that any user input query requires retrieving all historical dialogues, or being ineffective for all dialogues, meaning that any user input query does not require retrieving historical dialogues. Consequently, historical dialogues either interfere with normal intent recognition and processing, or fail to properly recognize and process user intents through historical dialogues, further causing inaccurate processing results from the large language model. Moreover, historical dialogues contain a large amount of invalid information, which can easily interfere with the processing of the large language model, further exacerbating the problem of inaccurate processing results from the large language model.
[0026] To address the shortcomings of related technologies, the embodiments in this specification, on the one hand, can flexibly acquire historical dialogues based on the completion status of intents by filtering historical dialogue information that have not completed intent recognition, thus avoiding the problems of acquiring all historical dialogues or acquiring none at all. On the other hand, by removing preset characters from historical dialogues that have not completed intent recognition and performing deduplication, invalid information in historical dialogues can be reduced, and the amount of data contained in historical dialogues can be decreased, thereby improving the accuracy of the model processing results and the efficiency of large language models in processing information.
[0027] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0028] Figure 2 This is a schematic diagram of the overall architecture of an information processing method provided in an embodiment of this specification. Figure 2As shown, the scheme may include terminal device 1, server 2, distributed cache 3, and large language model 4. Users can input queries on the target interface of terminal device 1; server 2 can obtain the user-input query and retrieve historical dialogues with incomplete intent recognition from the historical dialogue information contained in the target interface in the distributed cache 3; server 2 can preprocess the historical dialogues with incomplete intent recognition to obtain the target historical dialogue, and input the target historical dialogue and query into the large language model 4. The large language model 4 can process information based on the target historical dialogue and query, and feed the output results back to the server. In practical applications, server 2 may deploy distributed cache 3 and large language model 4; or server 2 may be able to call distributed cache 3 and large language model 4.
[0029] In such Figure 2 In the application scenarios shown, the server can connect to one or more terminal devices via a local area network (LAN), a wide area network (WAN), an internet connection, or other types of data networks. Figure 2 The servers mentioned can include, but are not limited to, any device, equipment, platform, or equipment cluster with computing and processing capabilities. Figure 2 The terminal devices in this context may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices.
[0030] Figure 3 This is a flowchart illustrating an information processing method provided in an embodiment of this specification.
[0031] From a programming perspective, the executor of the process can be a program hosted on an application server or application terminal. From a hardware perspective, the executor can be a server or an information processing system. It can be understood that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities.
[0032] like Figure 3 As shown, the process may include the following steps: Step 302: Obtain the query entered by the user on the target interface.
[0033] The target interface has corresponding historical dialogue information, which includes the user's historical queries and the historical answers of the large language model to the historical queries.
[0034] In the embodiments of this specification, the target interface can represent the interface through which the large language model in the terminal engages in dialogue with the user. A query can be an information request or operation command input by the user in the form of text, voice, image, etc., to meet their needs. A query can be information that the user wants the large language model to process. Historical dialogue information can be the dialogue information between the user and the large language model on the target interface prior to the user's input query.
[0035] In the embodiments described in this specification, historical dialogue information can be displayed on the target interface; alternatively, it can be hidden on the target interface, and the user can manipulate the controls to make the hidden historical dialogue information visible. Historical dialogue information can be stored in a distributed cache or in a pre-defined database for storing large language model data. Historical queries can be the user's historical input information on the target interface. Historical responses can be follow-up questions or answers from the large language model in response to historical queries. The large language model can be a multimodal large language model, capable of recognizing multimodal information such as images, text, and speech input by the user.
[0036] Step 304: Filter out historical dialogues that have not completed intent recognition from the historical dialogue information.
[0037] In the embodiments of this specification, a historical dialogue in which intent recognition is not completed can represent a dialogue in which the large language model does not provide a response that meets the user's expectations for the historical query input by the user. For example, if the historical query is "I want to buy supplementary teaching materials for my child," and the historical response of the large language model is "What subject's supplementary teaching materials do you want to buy for your child?", then it can be determined that the large language model did not provide a response that meets the user's expectations, and this historical query and historical response can be regarded as a historical dialogue in which intent recognition is not completed.
[0038] Step 306: Remove preset characters from the historical dialogues in which intent recognition was not completed to obtain preprocessed historical dialogues.
[0039] In the embodiments of this specification, the preset characters can be invalid information unrelated to the recognition intent of the large language model, such as punctuation marks ", ", ", etc., spaces, and interjections like "ah, ya, oh". Preset characters can be removed using regular expressions or other rules.
[0040] Step 308: Perform deduplication on the preprocessed historical dialogues to obtain the target historical dialogues.
[0041] In the embodiments described in this specification, deduplication can be achieved by removing historical dialogues with identical semantics or completely identical statements from the pre-processed historical dialogues. This avoids statements containing multiple repeated semantics and reduces the amount of data to be processed. In practical applications, the server can utilize a large model to remove invalid information and perform deduplication on historical dialogues; alternatively, it can process based on preset rules, such as regular expressions and hashing.
[0042] Step 310: Input the target historical dialogue and the query into the large language model to obtain the output result information of the large language model.
[0043] In the embodiments of this specification, the output result information can be the user intent successfully identified by the large language model after performing intent recognition on the target historical dialogue and query. User intent refers to the core needs or goals expressed by the user through input (such as text or voice), which can be specifically understood as "what problem the user wants to solve or what task the user wants to accomplish through this input". For example, if the user inputs "My meal card has no money, I want to recharge 500 yuan", then the user intent can be determined to be "recharge the meal card".
[0044] In the embodiments of this specification, the output result information may also be the processing result obtained by the large language model after identifying the user intent based on the target historical dialogue and query, and calling the business processing query that has a mapping relationship with the user intent.
[0045] While one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible execution order among many steps and does not represent the only possible execution order. The order of some steps may be adjusted according to actual needs, or some steps may be omitted. When the claims involve method steps, changes in the order of such steps, or parallel execution between steps, are also within the scope of protection of the claims.
[0046] Figure 3The method described above obtains the query entered by the user on the target interface, which may contain corresponding historical dialogue information. It filters out historical dialogues where intent recognition was incomplete, removes preset characters from these dialogues to obtain preprocessed historical dialogues, and then deduplicates them to obtain the target historical dialogue. The target historical dialogue and the query are then input into a large language model to obtain the output information. On the one hand, filtering out historical dialogues where intent recognition was incomplete avoids inputting historical dialogues where intent recognition was complete, improving the accuracy of the large language model in processing queries. On the other hand, removing preset characters from historical dialogues where intent recognition was incomplete reduces information irrelevant to the large language model's intent recognition, while also deduplicating the data input to the large language model, improving its efficiency in processing information.
[0047] based on Figure 3 In addition to the method described herein, this specification also provides some improved implementation methods, which will be described below.
[0048] In one or more embodiments of this specification, optionally, the step of filtering historical dialogues with incomplete intent identification from the historical dialogue information may specifically include: filtering historical dialogues with incomplete intent identification from the historical dialogue information according to a completion identifier; the completion identifier is used to identify historical dialogues with completed intent identification.
[0049] In the embodiments of this specification, a completion marker is a label used to identify the intent recognition status of historical dialogues. The completion marker allows the server to easily determine which dialogues in the historical dialogues have completed intent recognition and which have not. If multiple rounds of historical dialogues correspond to a single user intent and have completed intent recognition, then a completion marker can be assigned to these multiple rounds of historical dialogues. The completion marker can also be used to mark historical answers with preset key information and historical queries corresponding to those historical answers. Therefore, based on the completion marker, historical dialogues with incomplete intent recognition can be quickly filtered from the historical dialogue information, facilitating their use in subsequent information processing.
[0050] In practical applications, historical dialogues that have not completed intent recognition can be marked with an "incomplete" flag; or no flag can be added. The server can retrieve historical dialogue information from a distributed cache or pre-set database in chronological order based on the completion flag, identify the last round of dialogue among those marked with the completion flag, and determine other historical dialogues whose timestamps are after the original completion flag as those with incomplete intent recognition. Alternatively, the server can retrieve historical dialogue information from a distributed cache or pre-set database in chronological order based on the incomplete flag, and identify the last consecutive rounds or more marked with the incomplete flag as those with incomplete intent recognition.
[0051] In one or more embodiments of this specification, historical dialogues that have not completed intent recognition can also be filtered out using preset rules. Optionally, the step of filtering historical dialogues that have not completed intent recognition from the historical dialogue information may specifically include: identifying preset key information based on the information type of the historical responses; and filtering the historical dialogues that have not completed intent recognition based on the preset key information.
[0052] In the embodiments of this specification, the preset key information may differ for different information types. The preset key information may be in text, image, or audio format. The preset key information may be pre-set based on expert experience.
[0053] In practical applications, preset key information can be identified based on the information type of each historical response in the historical dialogue information. The historical responses can then be sequentially judged to determine whether they belong to historical dialogues where intent recognition has been completed, yielding the judgment result for each historical response. Alternatively, the judgment can be performed in parallel on each historical response in the historical dialogue information, yielding the judgment result for each historical response. Then, historical dialogues that have not completed intent recognition can be filtered out based on the judgment results of each historical dialogue and the dialogue time.
[0054] In practical applications, preset key information can be identified based on the information type of each historical response in the historical dialogue information. This allows for a sequential determination of whether a historical response represents a dialogue where intent recognition was incomplete, yielding a result for each response. Alternatively, the determination can be performed in parallel on each historical response, also yielding a result for each response. Then, based on the results of each historical dialogue and the dialogue time, dialogues with incomplete intent recognition can be filtered. This allows for the filtering of dialogues with incomplete intent recognition according to preset key information, even in the absence of completion markers, providing a valuable reference for large-scale query processing.
[0055] In one or more embodiments of this specification, optionally, the step of filtering the historical dialogues for which intent recognition has not been completed based on the preset key information may specifically include: if the preset key information exists, then the historical answers and the historical queries corresponding to the historical answers are determined as the historical dialogues for which intent recognition has not been completed.
[0056] In the embodiments of this specification, the preset key information may be information indicating that the user query has not been processed in the response output by the large language model. If the preset key information is contained in the historical responses, it can be indicated that the response content of the large language model to the historical query does not meet the user's expectations. Furthermore, it can be determined that the historical response and the corresponding historical query are historical dialogues with incomplete intent recognition.
[0057] In practical applications, historical dialogue information may contain multiple rounds of dialogue between the user and the large language model. If multiple rounds of historical dialogue exist, the pre-defined key information can be judged for each round of historical responses. If, after sorting by dialogue time from earliest to latest, the last historical response contains the pre-defined key information, then the last historical response and its corresponding historical query can be identified as historical dialogues with incomplete intent recognition. Alternatively, if multiple consecutive historical responses contain the pre-defined key information, then these multiple consecutive historical responses and their corresponding consecutive historical queries can be identified as historical dialogues with incomplete intent recognition. This allows for the identification of historical dialogues with incomplete intent recognition that can assist the large language model in processing queries.
[0058] In one or more embodiments of this specification, optionally, the step of filtering the historical dialogues that have not completed intent recognition based on the preset key information may specifically include: if the preset key information exists, determining the historical answers and the historical queries corresponding to the historical answers as historical dialogues that have completed intent recognition; and filtering the historical dialogues that have not completed intent recognition based on the historical dialogues that have completed intent recognition.
[0059] In the embodiments of this specification, the preset key information may be information indicating that the user query has been processed and completed in the response output by the large language model. If the preset key information is contained in the historical responses, it can be determined that the response content of the large language model to the historical query meets the user's expectations, and it can be determined that the historical response and the corresponding historical query belong to the historical dialogue for which intent recognition has been completed. Historical dialogues that do not belong to the historical dialogue for which intent recognition has been completed can be regarded as historical dialogues for which intent recognition has not been completed. Thus, historical dialogues for which intent recognition has been completed can be accurately identified through the above method, avoiding the use of historical dialogues for which intent recognition has been completed as historical dialogues for which intent recognition has not been completed, thereby improving the accuracy of information processing by the large language model.
[0060] In one or more embodiments of this specification, optionally, if the information type is text, the preset key information is a preset keyword, and the method may further include: determining whether the preset keyword exists in the historical answers.
[0061] In the embodiments of this specification, the preset keywords can be determined based on expert experience or based on the model's answering habits. For example, if the model habitually answers "Need to supplement XXX", then "supplement" can be used as the preset keyword. Symbols or words indicating questions, such as "?", "which", or "what", can also be used as preset keywords. The text type can indicate that there is no image-related information in the historical answers. Preset keyword information can be used to identify historical dialogues where intent recognition has not been completed, simply because of the specific content of the selected keyword. In practical applications, if other keywords are selected, historical dialogues where intent recognition has been completed can also be determined based on the preset keywords. For example, if the preset keyword is "complete", and the model answers "I have completed route planning, which can be divided into the following three routes...", then it can be indicated that the model's answer and the corresponding query are dialogues where intent recognition has been completed.
[0062] In the embodiments of this specification, the type can also be determined based on the user's intent. If the user's intent does not require the large model to output non-text results such as images or cards, the information type of the historical response can be determined to be text. For example, if the user's intent represents route information and the historical response is "Please provide destination information," then it can be determined that the preset keyword "provide" has been hit. Furthermore, it can be determined that the intent processing is incomplete. Thus, it is possible to determine whether the intent has been completely identified based on the historical responses of text type, improving the accuracy of filtering historical dialogues where the intent identification has not been completed.
[0063] In practical applications, if there are multiple rounds of historical dialogue in the historical dialogue information, and if the historical answers in a certain round of historical dialogue contain preset key information, then that certain round of historical dialogue can be identified as a historical dialogue in which intent recognition has not been completed; if the historical answers in the next round of historical dialogue do not contain preset key information, then that certain round of historical dialogue and the next round of historical dialogue can be identified as historical dialogues in which intent recognition has been completed. For example, in historical dialogue 1: historical query 1 is "I want to go to location B", and historical response 1 is "Where are you currently located?"; therefore, historical dialogue 1 can be determined as a historical dialogue where intent recognition was not completed. In historical dialogue 2: historical query 2 is "I am at the bus stop of shopping mall A in location C", and historical response 2 is "Do you want to travel by public transport?"; therefore, historical dialogue 2 can also be determined as a historical dialogue where intent recognition was not completed. In historical dialogue 3: historical query 3 is "Yes", and historical response 3 is "Okay, based on the query, we suggest you take bus number 25 from your current bus stop. Bus number 25 will arrive in 3 minutes"; therefore, historical dialogue 3 can be determined as a historical dialogue where intent recognition was completed. Thus, historical dialogues 1 and 2 can also be determined as historical dialogues where intent recognition was completed. Therefore, it is unnecessary to obtain historical dialogues 1, 2, and 3 as historical dialogues where intent recognition was not completed.
[0064] In one or more embodiments of this specification, optionally, if the information type is an image type, then the preset key information is image information, and the method may further include: determining whether the image information exists in the historical responses.
[0065] In the embodiments of this specification, the image type may include information such as images and cards. The information type of historical responses can be determined based on the content contained in the historical response information. If the historical response information contains image-related content, the information type of the historical response can be determined to be image type. For example, if the historical response information contains information such as "card ID" and "image ID", the information type of the historical response can be determined to be image type. If the historical response is image type and contains image information, the historical dialogue and corresponding historical query can be determined to be a historical dialogue for which intent recognition has been completed. If the historical response is image type and does not contain image information, the historical dialogue and corresponding historical query can be determined to be a historical dialogue for which intent recognition has not been completed.
[0066] In the embodiments of this specification, the information type of historical responses can also be determined based on user intent. For example, if the user intent is to query a certain card or generate an image, the information type of the historical responses can be determined to be image type. Taking cards as an example, if the card ID in the historical responses is empty, it can be determined that the corresponding card is not displayed in the historical responses, and the intent processing can be determined to be incomplete. Therefore, it is possible to determine whether the intent has been completely identified for historical responses of image type, improving the accuracy of filtering historical dialogues with incomplete intent recognition.
[0067] In one or more embodiments of this specification, hashing can be used to deduplicate historical dialogues that have not completed intent identification, thereby reducing duplicate information contained in these dialogues. Optionally, the preprocessed historical dialogues include a first preprocessed historical dialogue and a second preprocessed historical dialogue; the deduplication of the preprocessed historical dialogues may specifically include: hashing the first preprocessed historical dialogue to obtain a first hash value; hashing the second preprocessed historical dialogue to obtain a second hash value; if the first hash value and the second hash value are the same, then either the first preprocessed historical dialogue or the second preprocessed historical dialogue is removed.
[0068] In the embodiments of this specification, the first preprocessed historical dialogue or the second preprocessed historical dialogue can be any round of historical dialogue in the preprocessed dialogue. A round of historical dialogue can include a historical query and a historical answer corresponding to the historical query. Hash processing can refer to converting the preprocessed historical dialogue into a fixed-length string using a preset hash algorithm. If the first hash value and the second hash value are the same, it means that the core content of the first preprocessed historical dialogue and the second preprocessed historical dialogue are completely consistent, belonging to a duplicate dialogue. In this case, either preprocessed historical dialogue can be removed to avoid information redundancy while preserving data validity. If the first hash value and the second hash value are different, it means that the core content of the first preprocessed historical dialogue and the second preprocessed historical dialogue are different, not belonging to a duplicate dialogue. In this case, both the first preprocessed historical dialogue and the second preprocessed historical dialogue can be retained, thereby avoiding the deletion of valid information.
[0069] In one or more embodiments of this specification, to avoid semantically similar preprocessed historical dialogues, the preprocessed historical dialogues can be processed again. Optionally, before deduplicating the preprocessed historical dialogues, the process may further include: performing word segmentation on the first preprocessed historical dialogue to obtain a first word segmentation result; performing word segmentation on the second preprocessed historical dialogue to obtain a second word segmentation result; if the first word segmentation result and the second word segmentation result contain word segments with the same semantics, then replacing the word segments with the same semantics with word segments containing the same characters.
[0070] In the embodiments of this specification, word segmentation can refer to the process of splitting preprocessed historical dialogue into independent characters or words according to linguistic logic. Word segmentation is used to decompose preprocessed historical dialogue into the smallest semantic units that can be analyzed independently. Linguistic logic can represent grammatical rules and word boundaries in Chinese, etc. Word segmentation can be implemented using word segmentation tools, such as jieba, HanLP, and THULAC. The granularity of word segmentation can be determined based on actual needs. If fine-grained word segmentation is required, "book a hotel" can be split into two words: "book" and "hotel"; if coarse-grained word segmentation is required, "book a hotel" can be treated as a single word segmentation result. The first word segmentation result can contain one or more words; the second word segmentation result can also contain one or more words.
[0071] In this embodiment, semantic similarity calculation can be performed between any segment in the first segmentation result and each segment in the second segmentation result to determine the segment with the same semantic meaning in the first and second segmentation results. Alternatively, a semantic dictionary can be preset, and the segment with the same semantic meaning in the first and second segmentation results can be determined based on the semantic dictionary. For example, if the semantic dictionary contains "tomorrow": ["tomorrow", "next day", "tomorrow"], if any two of the segments "tomorrow", "tomorrow", "next day", and "tomorrow" exist, they can be determined to be segment with the same semantic meaning and can be uniformly replaced with "tomorrow". Replacing segment with the same semantic meaning with segment with the same character can be done by replacing one of the two identical segment with the other, or by using a target word with the same semantic meaning as the two identical segment and replacing both segment with the target word. This avoids generating different hash values for historical dialogues with the same semantic meaning during hash processing, improving the accuracy of deduplication.
[0072] In one or more embodiments of this specification, if the output result information is supplementary prompt information, it can be determined that the user intent is incomplete, and an incomplete identifier can be marked; the supplementary prompt information is used to prompt the user to supplement the information required for processing the business. If the output result information is a user intent, the business processing component that has a mapping relationship with the user intent can be invoked, and the user intent can be processed based on the user query and the historical dialogue of the incomplete intent identification to obtain the processing result information.
[0073] In practical applications, after recognizing the user intent, the large language model can also invoke the corresponding business logic to process the user intent, generate the corresponding processing result, and feed it back to the server. If the large language model cannot recognize the user intent, it can generate a model answer for follow-up questions; or, if the large language model can successfully recognize the user intent but lacks key information in processing the user intent, it can also generate a model answer for follow-up questions; the model answer for follow-up questions can be marked with an incomplete indicator.
[0074] In the embodiments of this specification, the output information of the large language model can represent the model's answer to the query. If the information type of the output information is image, it can be determined whether the output information contains image information. If the output information does not contain image information, it can be marked as incomplete; if the output information contains image information, it can be marked as complete. If the information type of the output information is text, it can be determined whether the output information contains preset keywords. If the output information contains preset keywords, it can be marked as incomplete; if the output information does not contain preset keywords, it can be marked as complete.
[0075] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they have not been described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.
[0076] According to the above explanation, Figure 4 This specification provides an overall flowchart of an information processing method, as illustrated in the embodiments below. Figure 4 As shown, it includes: Step 402: Obtain the query entered by the user on the target interface.
[0077] In the embodiments of this specification, the terminal device may retain multiple dialogue interfaces for the user to interact with the large language model, and the target interface may be the dialogue interface currently displayed on the terminal device. The terminal device may include controls for the user to switch between dialogue interfaces. Alternatively, the terminal device may contain multiple dialogue interfaces, and the target interface may be any of the multiple dialogue interfaces.
[0078] Step 404: Based on the completion marker, filter out the historical dialogues with incomplete intent recognition from the historical dialogue information.
[0079] In the embodiments described in this specification, historical dialogues with incomplete intent recognition can be retrieved from a distributed cache. In addition to the specific implementation methods described above, the server can also determine whether the query contains preset words. If the query contains preset words, it can be determined that the last group of historical dialogues marked with a completion flag also belongs to the historical dialogues with incomplete intent recognition. For example, in the last round of historical dialogue, the large language model provides a clear answer that meets preset requirements. The last round of historical dialogue, along with related historical dialogues, is marked with a completion flag. If the query is "Your answer is not what I want, I want XXX," it can be determined that the last round of historical dialogue and related historical dialogues still belong to the historical dialogues with incomplete intent recognition. A group of historical dialogues marked with a completion flag can contain one or more rounds of historical dialogues. In practical applications, multiple rounds of historical dialogues used to complete the same intent recognition can correspond to a single completion flag.
[0080] Step 406: Remove preset characters from the historical dialogues where intent recognition was not completed, and obtain the first preprocessed historical dialogue and the second preprocessed historical dialogue.
[0081] In practical applications, key data can be extracted from historical dialogues where intent recognition was not completed to obtain a first preprocessed historical dialogue and a second preprocessed historical dialogue. For example, if the historical query is "I want to go to location B" and the historical response is "Where are you located?", then the preprocessed historical dialogue can be obtained through key data extraction: historical query "go to location B"; historical response "Where are you located?". This is just an illustration, and the specific key data extraction results can be determined based on actual needs.
[0082] Step 408: Perform word segmentation on the first preprocessed historical dialogue and the second preprocessed historical dialogue respectively to obtain the first word segmentation result and the second word segmentation result.
[0083] Step 410: Replace semantically identical word segments in the first word segmentation result and the second analysis result with words containing the same characters.
[0084] Step 412: Determine if the hash values of the first preprocessed historical dialogue and the second preprocessed historical dialogue are the same. If yes, proceed to step 416: Remove either the first or second preprocessed historical dialogue to obtain the target historical dialogue. If no, proceed to step 414: Retain both the first and second preprocessed historical dialogues to obtain the target historical dialogue.
[0085] Step 418: Input the target historical dialogue and query into the large language model to obtain the output results.
[0086] Step 420: Determine whether the output result information meets the preset requirements. If yes, proceed to step 422: Store the output result information and the query in the distributed cache. If no, proceed to step 424: Store the output result information and the query completion flag in the distributed cache.
[0087] In the embodiments of this specification, if the output result information is text, the preset requirement is that the output result information contains a preset keyword; if the output result information is an image, the preset requirement is that the output result information does not contain image information. In addition to judging based on the above preset requirements, determining whether the output result has completed intent recognition can also be done using a large model to determine the judgment result.
[0088] The above methods achieve two main benefits. First, by using completion markers or preset rules to filter out historical dialogues that have not completed intent recognition, it is possible to flexibly acquire historical dialogues for input into a large language model to process queries, thereby improving the processing accuracy of the large language model. Second, by removing invalid information and deduplicating historical dialogues that have not completed intent recognition, the effectiveness of the target historical dialogues can be improved, while also reducing the amount of data in the target historical dialogues, thus improving the accuracy and efficiency of the model in processing information.
[0089] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods.
[0090] Figure 5 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of an information processing device.
[0091] like Figure 5 As shown, the device may include: The query acquisition module 502 is used to acquire the query entered by the user on the target interface; the target interface has corresponding historical dialogue information, which includes the user's historical queries and the historical answers of the large language model to the historical queries; The filtering module 504 is used to filter historical dialogues from the historical dialogue information to obtain those with incomplete intent recognition. The first preprocessing module 506 is used to remove preset characters from the historical dialogue in which intent recognition was not completed, and obtain the preprocessed historical dialogue. The second preprocessing module 508 is used to perform deduplication on the preprocessed historical dialogue to obtain the target historical dialogue. The model processing module 510 is used to input the target historical dialogue and the query into the large language model to obtain the output result information of the large language model.
[0092] based on Figure 5 The embodiments of this specification also provide some specific implementation schemes of the method, which are described below.
[0093] Optionally, the filtering module can be specifically used to: filter historical dialogues with incomplete intent recognition from the historical dialogue information according to the completion identifier; the completion identifier is used to identify historical dialogues with completed intent recognition.
[0094] Optionally, the filtering module can be specifically used to: identify preset key information based on the information type of the historical answers; and filter the historical dialogues in which intent recognition has not been completed based on the preset key information.
[0095] Optionally, the filtering module can be specifically used to: if the preset key information exists, determine the historical answers and the historical queries corresponding to the historical answers as the historical dialogues for which intent recognition has not been completed.
[0096] Optionally, the filtering module can be specifically used to: if the preset key information exists, determine the historical answers and the historical queries corresponding to the historical answers as historical dialogues that have completed intent recognition; and filter the historical dialogues that have not completed intent recognition based on the historical dialogues that have completed intent recognition.
[0097] Optionally, the filtering module can be specifically used to: determine whether the preset keyword exists in the historical answers.
[0098] Optionally, the filtering module can be specifically used to: determine whether the image information exists in the historical answers.
[0099] Optionally, the second preprocessing module may be specifically used to: perform hash processing on the first preprocessed historical dialogue to obtain a first hash value; perform hash processing on the second preprocessed historical dialogue to obtain a second hash value; if the first hash value and the second hash value are the same, then remove any one of the first preprocessed historical dialogues and the second preprocessed historical dialogues.
[0100] Optionally, the second preprocessing module may be specifically used to: perform word segmentation on the first preprocessed historical dialogue to obtain a first word segmentation result; perform word segmentation on the second preprocessed historical dialogue to obtain a second word segmentation result; if the first word segmentation result and the second word segmentation result contain word segments with the same semantics, then replace the word segments with the same semantics with word segments with the same characters.
[0101] It is understood that the modules mentioned above refer to computer programs or program segments used to perform one or more specific functions. Furthermore, the distinction between these modules does not imply that the actual program code must also be separate.
[0102] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0103] The above is an illustrative scheme of an information processing device according to this embodiment. It should be noted that the technical solution of this information processing device and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the information processing device, please refer to the description of the technical solution of the information processing method described above.
[0104] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.
[0105] Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification.
[0106] The computing device 600 includes: Memory 610 and processor 620; The memory 610 is used to store computer programs / instructions, and the processor 620 is used to execute the computer programs / instructions, which, when executed by the processor 620, implement the steps of the information processing method.
[0107] Specifically, the components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and the database 650 is used to store data.
[0108] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0109] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.
[0110] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.
[0111] The processor 620 executes the computer instructions to implement the steps of the information processing method.
[0112] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the information processing method described above.
[0113] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the steps of the information processing method as described above.
[0114] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the information processing method described above.
[0115] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described information processing method.
[0116] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the information processing method described above.
[0117] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the embodiments of apparatus, devices, media, and products, since they are basically similar to the method embodiments, the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the method embodiments. The apparatus, devices, media, and products provided in the embodiments of this specification correspond to the methods; therefore, the apparatus, devices, media, and products also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding apparatus, devices, media, and products will not be repeated here.
[0118] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0119] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0120] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0121] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0122] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0123] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, the invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0124] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0125] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0126] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0127] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0128] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0129] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital character versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0130] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0131] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. An information processing method, comprising: obtaining a query input by a user at a target interface; the target interface having corresponding historical dialogue information, the historical dialogue information including historical queries of the user and historical answers of a large language model to the historical queries; screening historical dialogues with uncompleted intent recognition from the historical dialogue information; removing preset characters in the historical dialogues with uncompleted intent recognition to obtain preprocessed historical dialogues; de-duplicating the preprocessed historical dialogues to obtain target historical dialogues; inputting the target historical dialogues and the query into the large language model to obtain output result information of the large language model.
2. The method of claim 1, wherein the screening of the historical dialogues with uncompleted intent recognition from the historical dialogue information specifically comprises: screening the historical dialogues with uncompleted intent recognition from the historical dialogue information according to a completion identifier; the completion identifier being used to identify historical dialogues with completed intent recognition.
3. The method of claim 1, wherein the screening of the historical dialogues with uncompleted intent recognition from the historical dialogue information specifically comprises: identifying preset key information based on an information type of the historical answer; screening the historical dialogues with uncompleted intent recognition based on the preset key information.
4. The method of claim 3, wherein the screening of the historical dialogues with uncompleted intent recognition based on the preset key information specifically comprises: if the preset key information exists, determining the historical answer and a historical query corresponding to the historical answer as the historical dialogues with uncompleted intent recognition.
5. The method of claim 3, wherein the screening of the historical dialogues with uncompleted intent recognition based on the preset key information specifically comprises: if the preset key information exists, determining the historical answer and a historical query corresponding to the historical answer as historical dialogues with completed intent recognition; screening the historical dialogues with uncompleted intent recognition based on the historical dialogues with completed intent recognition.
6. The method of claim 4, wherein if the information type is a text type, the preset key information is a preset keyword, and the method further comprises: determining whether the preset keyword exists in the historical answer.
7. The method of claim 5, wherein if the information type is an image type, the preset key information is image information, and the method further comprises: determining whether the image information exists in the historical answer.
8. The method of claim 1, wherein the preprocessed historical dialogues include first preprocessed historical dialogues and second preprocessed historical dialogues, and the de-duplicating of the preprocessed historical dialogues specifically comprises: hashing the first preprocessed historical dialogues to obtain a first hash value; hashing the second preprocessed historical dialogues to obtain a second hash value; if the first hash value is the same as the second hash value, removing any one of the first preprocessed historical dialogues and the second preprocessed historical dialogues.
9. The method of claim 8, before the deduplication processing on the preprocessed historical dialogue, further comprising: performing word segmentation processing on the first preprocessed historical dialogue to obtain a first word segmentation result; performing word segmentation processing on the second preprocessed historical dialogue to obtain a second word segmentation result; if the first word segmentation result and the second word segmentation result contain a same semantic word segmentation, replacing the same semantic word segmentation with a character same word segmentation.
10. An information processing apparatus, comprising: a query obtaining module configured to obtain a query input by a user on a target interface; the target interface has corresponding historical dialogue information, and the historical dialogue information includes historical queries of the user and historical answers of a large language model to the historical queries; a screening module configured to screen historical dialogues with incomplete intent recognition from the historical dialogue information; a first preprocessing module configured to remove preset characters in the historical dialogues with incomplete intent recognition to obtain preprocessed historical dialogues; a second preprocessing module configured to perform deduplication processing on the preprocessed historical dialogues to obtain target historical dialogues; a model processing module configured to input the target historical dialogues and the query into the large language model to obtain output result information of the large language model.
11. A computing device, comprising: a memory and a processor; the memory is configured to store a computer program or instructions, and the processor is configured to execute the computer program or instructions, and the computer program or instructions, when executed by the processor, implement the steps of the method in any one of claims 1 to 9.