Information processing method and device, equipment, storage medium and program product

By building a collaborative knowledge base for operations and maintenance and machine learning models, the limitations of existing operations and maintenance question-and-answer systems in handling complex operations and maintenance issues have been overcome, resulting in more efficient and accurate operations and maintenance question-and-answer, and improved operational efficiency and user experience.

CN121614568APending Publication Date: 2026-03-06JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511803325.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Existing operational Q&A technologies for vertical operations and maintenance have limitations when dealing with complex and professional operations and maintenance issues. These limitations include a lack of understanding of professional terminology and scenario logic in the operations and maintenance field, high costs of knowledge graph construction and maintenance, insufficient multimodal data processing capabilities, low efficiency in long text processing, and insufficient rejection models. As a result, the accuracy and practicality of operations and maintenance scenario Q&A are insufficient.

Method used

By building a specialized knowledge base for operations and maintenance, and using machine learning models to determine the weights of text fragments, responses to user input are generated. This involves the collaborative work of rejection models, intent recognition models, recommendation models, text summarization models, and language models, enabling accurate processing of user input and proactive question answering.

Benefits of technology

It improves the accuracy and efficiency of question answering in operation and maintenance scenarios, reduces the false rejection rate, enhances the ability to process long texts, strengthens the fusion analysis of multimodal data, reduces the latency of knowledge base updates, and improves user experience and the professional credibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121614568A_ABST
    Figure CN121614568A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to an information processing method and device, equipment, a storage medium and a program product. The method comprises the following steps: in response to receiving user input, determining a query indicated by the user input; obtaining at least one text fragment associated with the query from the knowledge base; determining the weight of each text fragment in the at least one text fragment based on the relevance between each text fragment in the at least one text fragment and the field corresponding to the query and the timeliness of each text fragment; and generating a reply for the user input based on the user input, the at least one text fragment and the weight of each text fragment by using a first machine learning model. Therefore, the accuracy of the generated reply can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computer technology, and more specifically, to methods, apparatus, electronic devices, computer-readable storage media, and computer program products for information processing. Background Technology

[0002] In the field of Artificial Intelligence for IT Operations (AIOps), vertical-specific operational question-answering systems are crucial tools for improving operational efficiency and accuracy. Current operational question-answering technologies primarily rely on general question-answering models (such as Transformer-based pre-trained language models), rule-based question-answering systems, and traditional natural language processing techniques. These methods have limitations when handling complex operational scenarios and diverse operational needs, thus necessitating an improved question-answering solution specifically designed for the operational domain. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for information processing is provided. The method includes: in response to receiving user input, determining a query indicated by the user input; retrieving at least one text fragment associated with the query from a knowledge base; determining weights of each text fragment in the at least one text fragment based on the relevance between each text fragment and the domain corresponding to the query, and the timeliness of each text fragment; and generating a response to the user input using a first machine learning model, based on the user input, the at least one text fragment, and the weights of each text fragment.

[0004] In a second aspect of this disclosure, an apparatus for information processing is provided. The apparatus includes: a query determination module configured to determine a query indicated by user input in response to receiving user input; a text acquisition module configured to acquire at least one text fragment associated with the query from a knowledge base; a weight determination module configured to determine the weights of each text fragment in the at least one text fragment based on the relevance between each text fragment and the domain corresponding to the query, and the timeliness of each text fragment; and a response generation module configured to generate a response to the user input using a first machine learning model, based on the user input, the at least one text fragment, and the weights of each text fragment.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The medium stores a computer program that, when executed by a processor, implements the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided. The product includes a computer program, wherein when executed by a processor, the computer program implements the method according to a first aspect of this disclosure.

[0008] It should be understood that the description in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of various implementations of this disclosure will become more apparent in the following detailed description, taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown; Figure 2 An example architecture of a system for information processing according to some embodiments of the present disclosure is shown; Figure 3 An example architecture for constructing a knowledge base example process according to some embodiments of this disclosure is shown; Figure 4 A flowchart of a method for information processing according to some embodiments of the present disclosure is shown; Figure 5 A schematic structural block diagram of an apparatus for information processing according to some embodiments of the present disclosure is shown; and Figure 6 A block diagram of a computing device in which embodiments of the present disclosure may be implemented is shown. Detailed Implementation

[0010] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0011] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0012] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0013] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0014] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and authorization should be obtained from the relevant users. Among them, relevant users may include any type of rights holder, such as individuals, enterprises, and groups.

[0015] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly inform the user that the requested operation will require obtaining and using the user's information, thereby enabling the relevant user to choose whether to provide information to the software or hardware such as the electronic device, application, server, or storage medium that performs the operation of the technical solution disclosed herein based on the prompt message.

[0016] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide information to the electronic device.

[0017] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0018] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0019] A neural network is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs. It typically consists of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the inputs to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the layer above.

[0020] As mentioned earlier, the question-and-answer solutions commonly used in the field of intelligent operation and maintenance have certain limitations when dealing with complex operation and maintenance scenarios and diverse operation and maintenance needs, mainly in the following aspects.

[0021] First, while question-answering systems based on general question-answering models perform well in natural language processing tasks, they lack targeted optimization for O&M-specific terminology, scenario logic, and knowledge structures in O&M vertical scenarios. Therefore, general question-answering models lack a deep understanding of O&M-specific knowledge and terminology. When dealing with highly specialized and logically complex O&M problems, these models often struggle to provide accurate answers. For example, when answering specialized questions such as database performance tuning and root cause analysis, general models fail to accurately understand the complex context of the O&M scenario, leading to a significant decrease in accuracy.

[0022] Second, for question-and-answer systems based on rule engines, although they can handle some common operational issues, the coverage of rules is limited for complex operational scenarios and unknown operational patterns, making it difficult to adapt to rapidly changing business needs and complex operational environments.

[0023] Third, for question-answering systems based on traditional natural language processing technologies, these technologies (such as keyword matching and template matching) have problems such as insufficient understanding and poor contextual relevance when dealing with complex queries in the field of operations and maintenance, making it difficult to meet the high accuracy requirements of operations and maintenance question-answering.

[0024] Fourth, the construction and maintenance costs of knowledge graphs in the field of intelligent operations and maintenance (O&M) are currently high. While knowledge graphs can provide structured knowledge representations, their construction and maintenance in the O&M field require significant professional knowledge and manpower, and their update speed struggles to keep pace with the rapidly changing O&M environment. Furthermore, the chunking strategies employed using general natural language processing (NLP) technologies (such as fixed-length segmentation or simple semantic segmentation) are ineffective in handling the special structures in O&M technical documents (such as error code tables and alarm rule trees), leading to the loss of crucial information during knowledge retrieval. For example, splitting an O&M manual containing multiple related parameters into independent paragraphs can disrupt the logical connections between parameters. In addition, O&M knowledge bases are updated frequently (e.g., weekly vulnerability patches and monthly configuration changes), but current systems often use periodic full updates, resulting in several days of delay in responding to hot issues (such as solutions for sudden failures). In time-sensitive scenarios such as emergency response to CVE (Common Vulnerabilities & Exposures) vulnerabilities, this lag can pose serious O&M risks.

[0025] Fifth, current question-and-answer systems lack the ability to process multimodal data. Operational scenarios often involve multiple types of data (such as logs, monitoring data, configuration information, etc.). Current systems lack effective fusion and comprehensive analysis capabilities when processing this multimodal data, making it difficult to provide comprehensive operational solutions.

[0026] Sixth, current systems typically employ a single retrieval model (such as keyword or vector-based retrieval), which struggles to balance precise matching (e.g., error codes) and semantic generalization (e.g., "service unavailable") for specific operational scenarios, resulting in a very low average recall rate. Furthermore, the system is weak in handling long-tail issues and is easily affected by outdated or low-quality knowledge fragments. Additionally, traditional re-ranking models lack the fine-grained matching capability for operational domain-specific issues (e.g., "SSL certificate chain verification failed"), requiring secondary manual screening.

[0027] Seventh, the current system is inefficient at processing long texts. Large models are limited by token length, making it difficult to directly process long texts such as error logs. General-purpose summarization models have low accuracy in extracting key information in operational scenarios, leading to context truncation and semantic loss. Furthermore, the system lacks intelligent interaction; it lacks proactive recommendation capabilities after a user asks a question, and general-purpose recommendation models, due to a lack of understanding of operational context (such as work order history and frequently asked questions), have low recommendation hit rates, requiring users to repeatedly filter recommendations.

[0028] Eighth, current systems generally lack rejection models specifically designed for the operations and maintenance (O&M) domain. When users ask irrelevant, casual questions or questions outside the scope of the knowledge base (such as "How do I restart the server?" and "What's the weather like today?"), the model still forces the generation of answers in the tone of an O&M expert. This "irrelevant answer" phenomenon not only degrades the user experience but also damages the system's professional credibility.

[0029] In summary, current operational Q&A technologies for vertical operations and maintenance (O&M) have many limitations when dealing with complex and professional O&M issues. There is an urgent need for an improved Q&A system for the O&M field to enhance the accuracy and practicality of O&M Q&A in O&M scenarios.

[0030] In view of this, embodiments of the present disclosure provide an improved solution for information processing to at least solve the above-mentioned problem. In this solution, in response to receiving user input, the query indicated by the user input is determined; at least one text fragment associated with the query is obtained from a knowledge base; the weights of each text fragment in the at least one text fragment are determined based on the relevance between each text fragment and the domain corresponding to the query, and the timeliness of each text fragment; and a response to the user input is generated using a first machine learning model based on the user input, the at least one text fragment, and the weights of each text fragment. Thus, by determining the weights of each text fragment by considering the relevance between the text fragment obtained from the knowledge base and the domain corresponding to the query, and the timeliness of the text fragment, the model can generate a response to the user input based on text fragments more relevant to the query indicated by the user input, thereby improving the accuracy of the response.

[0031] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, a target application 112 is installed on a client device 110. A user 130 can interact with the application service component target application 112 via the client device 110 and / or an attached device to the client device 110.

[0032] In some embodiments, the target application 112 may be downloaded and installed on the client device 110. In some embodiments, the target application 112 may also be accessed in other ways, such as via a webpage. In some embodiments, in Figure 1 In environment 100, in response to the launch of target application 112, client device 110 can present interface 140 of target application 112. Interface 140 can be, for example, the interactive interface of target application 112.

[0033] In some embodiments, the target application 112 may have intelligent dialogue and information processing capabilities. For example, the target application 112 may utilize a machine learning model to perform user question-and-answer. In embodiments of this disclosure, the target application 112 is used to interact with the user 130 to assist the user 130 in using a terminal or processing and searching information. In some embodiments, an interaction window with the target application 112 may be presented in the interface 140. In the interaction window, the user 130 can engage in dialogue with the target application 112 by inputting natural language, images, audio files, video files, web page files, etc., to instruct the target application 112 to assist in completing various tasks, including operations on content entities; or to instruct the target application 112 to complete question-and-answer, query, or search. In some embodiments, the interaction between the user 130 and the target application 112 may include interaction between the user 130 and a digital assistant. For example, the user's input to the target application 112 may be issued to a machine learning model, and the response from the target application 112 to the user 130 may be issued by the machine learning model to the user 130.

[0034] In some embodiments, client device 110 communicates with server 120 to provide services to target application 112. Client device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, client device 110 may also support any type of user-facing interface (such as "wearable" circuitry). Server 120 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc. Server 120 may deploy one or more various machine learning models 122 to provide services to target application 112. In some embodiments, model 122 can be a machine learning model, a deep learning model, a learning model, a neural network, etc. In some embodiments, the model may be based on a language model (LM). Language models can acquire question-answering capabilities by learning from large corpora. Model 122 can also be based on other appropriate models.

[0035] In some implementations, the implementation of at least some functions of the target application 112, and / or the implementation of at least some functions in the target application 112, may be based on models. During the creation or operation of the target application 112, one or more models 122 may be invoked, such as the capabilities of model 122. In the target application 112, model 122 may be used to understand user input and to provide responses to the user based on the output of model 122.

[0036] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0037] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0038] Figure 2 An example architecture of a system 200 for information processing according to some embodiments of the present disclosure is shown. For example... Figure 2As shown, the system 200 for information processing according to some embodiments can be implemented or included in client device 110 and / or server 120. In the following description of embodiments, for the purposes of discussion, the implementation at client device 110 is described as an example. Operations implemented at client device 110 can be implemented with the assistance of server 120 or by requesting server 120 and obtaining a response from server 120. In some embodiments, one or more models discussed below can be deployed locally or remotely on client device 110. Therefore, one or more models can be invoked directly by client device 110 or by client device 110 via a request to server 120, which in turn invokes them.

[0039] In some embodiments, client device 110 acquires user input 210 and provides the user input to rejection model 212. For example, target application 112 of client device 110 acquires user input 210 and provides the user input to rejection model 212. Rejection model 212 determines whether the user input is associated with a predetermined domain based on the user input. Figure 2 The architecture is designed for specific training and configuration to handle question-answering requests in a particular domain. In some embodiments, the predetermined domain is an operations domain. It is understood that in other embodiments, the predetermined domain can also be other domains, allowing the model to perform question-answering based on knowledge within that domain.

[0040] In response to the rejection model 212 determining that the user input is not related to the predetermined domain, the client device 110 or the target application 112 provides a prompt message to inform the user that the user input is not related to the predetermined domain. The prompt message may be, for example, "This question does not belong to the operations and maintenance domain; the system cannot answer it." In response to the rejection model 212 determining that the user input is related to the predetermined domain, the client device 110 or the target application 112 provides the user input 210 to the intent recognition model 216, recommendation model 214, or text summarization model 218, etc., for further processing.

[0041] In some embodiments, the rejection model 212 is implemented based on the RoBERTa (Robustly Optimized BERT Approach) architecture. For example, the rejection model 212 can be implemented by pre-training or fine-tuning a RoBERTa model. In some embodiments, the rejection model 212 can be obtained by training or fine-tuning a RoBERTa model using data from the operations and maintenance domain as positive examples and data from the general domain as negative examples.

[0042] In some embodiments, the rejection model 212 may output a correlation result between the user input 210 and a predetermined domain, along with a corresponding confidence score. The client device 110 or the target application 112 determines that the user input 210 is associated with the predetermined domain in response to the correlation result indicating that the user input 210 is associated with the predetermined domain and the confidence score is greater than a first threshold. Alternatively, the client device 110 or the target application 112 determines that the user input 210 is not associated with the predetermined domain in response to the correlation result indicating that the user input 210 is not associated with the predetermined domain and the confidence score is greater than the first threshold. In some embodiments, the first threshold is, for example, any suitable value such as 0.7 (70%), 0.8 (80%), or 0.9 (90%).

[0043] In some embodiments, the relevance outcome can be expressed as a binary decision outcome such as relevance / non-relevance or acceptance / rejection. In some embodiments, the rejection model 212 can also support fine-grained relevance outcomes. For example, the relevance outcome can include multiple levels such as "completely irrelevant", "possibly irrelevant", "possibly relevant", "domain relevant but lacking in capability", and "domain relevant".

[0044] In some embodiments, in response to the relevance result output by the rejection model 212 indicating that the user input 210 is associated with a predetermined domain, but the confidence score is greater than a second threshold and less than a first threshold, the client device 110 or the target application 112 provides the user input 210 to the text summarization model 218. The text summarization model 218 extracts key information from the user input 210, such as key fields, and then provides the key information to the rejection model 212. The rejection model 212 further generates a relevance result and a confidence score based on the key information. In this way, by combining with the text summarization model 218, the target application 112 can reduce the false rejection rate or probability of questions in the predetermined domain and improve processing accuracy. As an example, if the user input 210 is "My pod keeps crashing back off", the rejection model 212 initially judges that the user input may be related to the predetermined domain, but the confidence score is 0.65 (the first threshold is 0.7). At this time, it triggers entry into the text summarization model 218. Text summarization model 218 extracts the key field `{"resource": "pod", "state": "CrashLoopBackOff"}` from user input 210. Rejection model 212 performs a secondary decision based on the key field, at which point the confidence score increases to 0.82. Client device 110 or target application 112 determines that user input 210 is associated with a predetermined domain and proceeds to the next step.

[0045] Once the client device 110 or the target application 112 determines that the user input 210 is associated with a predetermined domain based on the relevance results and confidence scores output by the rejection model 212, the client device 110 or the target application 112 provides the user input 210 to the recommendation model 214, the intent recognition model 216, the text summarization model 218, and the rewriting model 220, etc.

[0046] In some embodiments, recommendation model 214 generates a set of recommended queries based on user input 210 for the user to choose from. Target application 112 may, in response to the user's selection of a query from the set of recommended queries, use the user-selected query as the final query to be processed.

[0047] In some embodiments, the target application 112 acquires at least one of user historical behavior data and contextual information, and provides the acquired historical behavior data and contextual information to the recommendation model 214. The recommendation model 214 determines a set of recommended queries for the user to choose from based on the user input 210 and at least one of the user historical behavior data and contextual information. Historical behavior data may include, for example, the user's previous click data and dwell time on the recommended queries. Historical behavior data may also include historical work orders, user session records, etc. Contextual information may include system service status, on-duty engineer information, etc.

[0048] In some embodiments, the recommendation model 214 can be implemented based on a hybrid architecture consisting of an extreme gradient boosting (XGBoost) model and a transformer model. The recommendation model 214 can utilize XGBoost to process structured features (such as user job level and service cluster affiliation), and the Transformer model to encode the semantics of the query, calculating the similarity to questions in the knowledge base 224. During the training of the recommendation model 214, the click-through rate of user recommendations and the resolution rate of user questions can be used as optimization objectives to train the model, thereby achieving a higher recommendation hit rate and a significant improvement over models without domain context.

[0049] The training of recommendation model 214 can be completed through the following process. First, feature engineering is performed to extract user behavior features, question text features, and context features from the user's historical behavior, historical questions, and context. Then, synthetic model training is performed, using XGBoost to process structured features and Transformer to process text features for multi-task learning.

[0050] In some embodiments, the intent recognition model 216 can determine at least one processing intent indicated by user input 210 based on user input 210. The intent recognition model 216 can provide at least one processing intent to the routing policy engine 232 and the rewriting model 220. At least one processing intent can be identified by a structured intent label. In some embodiments, the intent recognition model 216 can generate a structured intent label based on user input 210 and context information. As an example, user input 210 is "What to do about MySQL slave database synchronization delay", and context information may include information such as historical sessions and service topology. The structured intent label output by the intent recognition model 216 is, for example, {"domain": "database", "task": "replication fault troubleshooting", "urgency": "high"}. The structured intent label indicates information such as the domain involved in the processing intent, the processing task, and the urgency level using a predetermined template.

[0051] In some embodiments, the intent recognition model 216 can be built based on any suitable model or architecture (e.g., a bidirectional encoder representation from transformers), and the model can be fine-tuned or pre-trained based on maintenance work order data. In some embodiments, the intent recognition model 216 can recognize up to 200 or more intent categories, including troubleshooting, configuration optimization, capacity planning, etc.

[0052] In some embodiments, an operation and maintenance intent graph can also be constructed. The intent recognition model 216 can automatically invoke tools (such as directly triggering "large model SQL diagnosis" after recommendation) based on the operation and maintenance intent graph to handle the user's problem. The construction of the operation and maintenance intent graph can be completed through the following steps. First, a data source is formed based on historical work orders, user session records, popular problem databases, etc., and then a graph neural network is used to mine the correlation paths between problems from the data source. The construction of the operation and maintenance intent graph can be used to realize the semantic association of "problem-solution-toolchain" (such as alarm → root cause → automatic triggering of diagnostic tools). As an example, if the recommended query generated by the user input 210 or the recommendation model 214 selected by the user is "MySQL query is slow", the intent recognition model 216 can determine the cause as an index building problem based on the constructed operation and maintenance intent graph, the operation to be performed is index optimization, and the tool to be used is an SQL diagnostic tool. This forms the semantic association of "problem-solution-toolchain", which can realize proactive question answering based on context-based dynamic recommendation. The information processing system of this disclosure can support automatic tool invocation, achieve a high recommendation hit rate through the optimized recommendation model 214, and form a "question-answer-recommendation-execution" closed loop through collaboration with the intent recognition model 216 to achieve proactive question answering and problem solving.

[0053] In some embodiments, the routing policy engine 232 can generate routing policies based on the processing intent generated by the intent recognition model 216 and the real-time system status. The routing policy indicates the routing path for the processing intent indicated by the user input. The real-time system status may include information such as current load and expert online status. Routing policies can be represented using structured tags. As an example, a routing policy could be `{"target": "SQL diagnostic tool", "priority": "P0", "handler": "automatic execution"}`. Routing policies can pre-define templates indicating the invocation tool, processing priority, processing method, etc. It should be noted that routing policies can include various types, such as automated routing policies, manual routing policies, and hybrid routing policies. Automated routing policies target simple problems, which can directly trigger the toolchain for processing (e.g., "display disk space" → invoking `df -h`). Manual routing policies target complex problems, delegating them to the corresponding domain expert (e.g., "kernel parameter optimization" → Linux system group). Hybrid routing policies target scenarios requiring human-machine collaboration (e.g., "canary release anomaly" → first invoking log analysis tools, then delegating to operations engineers). It should be understood that the routing policy engine 232 can also generate routing policies based on the knowledge of the knowledge base 224. It is also understood that the routing policy generated by the routing policy engine 232 can be provided to the language model 230 along with the user input 210, and the language model 230 can generate a response to the user input 210 based on the user input 210 and the routing policy.

[0054] In the embodiments of this disclosure, proactive question answering can be achieved through the collaboration of recommendation model 214, intent recognition model 216 and routing policy engine 232, which greatly improves the processing efficiency of the system.

[0055] Continue to refer to Figure 2 In some embodiments, the text summarization model 218 can extract structured information from user input 210 and generate a standardized summary for user input 210 based on the structured information and a summarization template. In some embodiments, the text summarization model 218 can be a lightweight summarizer built based on an operations and maintenance variant model, and the text summarization model 218 has a relatively small number of parameters. In some embodiments, the text summarization model 218 can extract structured information from long texts by pre-training and learning key fields of logs (such as error codes, timestamps, and service names) to generate standardized summaries.

[0056] The training of the text summarization model 218 can be completed through the following process: First, pre-training is performed on an operations and maintenance log dataset based on the LogBERT architecture to learn the representation of key fields in the logs; second, the model is trained using a manually annotated standardized summary template as supervised data, enabling the model to output standardized summaries. In the data preparation stage, a standardized summary template is generated through manual annotation. This template defines the fields or information that the summary must include, such as defining a JSON schema to include required fields. During the model training phase, the summary template serves as supervised data to guide LogBERT in learning how to extract structured information from raw logs. In the model inference phase, the summary template can be used to constrain the model's output format, ensuring that the generated summary meets the needs of operations and maintenance analysis.

[0057] As an example, user input 210 includes the log message "2023-10-05 08:22:15 ERROR[Broker-1] Failed to process request from producer clientId=producer-3, error=NotEnoughReplicasException: Message queue is full......", and the generated normalized summary is, for example, {"timestamp": "2023-10-05T08:22:15","service": "Kafka Broker-1","error_type":"NotEnoughReplicasException","root_cause": "Message queue full","affected_client": "producer-3"}.

[0058] In some embodiments, in response to determining that user input 210 includes log information, target application 112 provides user input 210 or the log information included in user input 210 to text summarization model 218. Text summarization model 218 extracts structured information from the log information and generates a standardized summary for the user input based on the structured information. After generating the standardized summary, text summarization model 218 can provide the standardized summary to rewriting model 220, dual-path retrieval model 222, or language model 230. Rewriting model 220 can rewrite the standardized summary to add contextual information. Dual-path retrieval model 222 can retrieve relevant knowledge from knowledge base 224 based on the standardized summary, for example, by retrieving relevant knowledge fragments in knowledge base 224 using fields included in the standardized summary as keywords. Language model 230 can generate a response for the user input based on the standardized summary.

[0059] In this embodiment, key information from logs is extracted using a text summarization model. This reduces the token input to a language model (e.g., LLM) by summarizing the data, thereby improving the reasoning efficiency of large models in knowledge-based question answering within operational scenarios. It also addresses the issue of large models reaching their token input limit, which prevents question answering. More specifically, because structured summarization increases the density of key information, the language model (e.g., LLM) only needs to process high-information-density summaries, significantly improving reasoning speed while avoiding token overrun issues caused by excessively long original logs. Furthermore, the pre-training of the text summarization model enables it to achieve higher recognition accuracy for key fields in operational scenarios (such as error codes and service names), compensating for the shortcomings of general-purpose models in vertical scenarios. The collaboration between the text summarization model and the language model achieves a dual improvement in efficiency and accuracy.

[0060] This embodiment of the disclosure significantly reduces token consumption for raw log input by combining a text summarization model and a language model, thereby greatly improving the inference speed of the language model and enabling second-level analysis of thousands of words of logs. Furthermore, the accuracy of extracting key log information (such as error root causes and service impacts) is greatly increased, significantly outperforming general summarization models and breaking through the efficiency bottleneck of models processing long texts.

[0061] Continue to refer to Figure 2 The rewrite model 220 can rewrite user input 210 to obtain enhanced user input. The changes to the enhanced user input relative to user input 210 can include using standardized terminology, including more operational context information, and associating Chinese error messages with English log keywords. For example, the user input "The container is down" (in colloquial terms) is rewritten as "The Pod is in CrashLoopBackOff state." The user input "API response is slow" is rewritten by automatically adding the service name and time period, becoming "The order service API is delayed during peak evening hours."

[0062] In some embodiments, rewriting model 220 can generate at least one candidate user input based on user input 210 and at least one processing intent generated by intent recognition model 216. In some embodiments, when rewriting model 220 generates more than two candidate user inputs, recommendation model 214 selects one of the candidate user inputs as the user input based on the user service topology. Then, rewriting model 220 generates an enhanced user input based on the selected candidate user input and contextual information related to user input 210. As an example, user input 210 is, for example, "My database is lagging." Intent recognition model 216 determines that the processing intent is related to database performance based on user input 210, but the expression is somewhat ambiguous. At this time, rewriting model 220 generates candidate rewrites based on user input 210: "MySQL database query performance is degraded," "Redis cache response latency is increased," and "MongoDB write throughput is reduced." Recommendation model 214 selects "MySQL database query performance is degraded" as the user input based on the user service topology (which is known to use MySQL). Then, Model 220 was rewritten to further enhance the "MySQL database query performance degradation" based on context, resulting in "the average latency of SELECT queries in MySQL databases increased from 50ms to 800ms between 08:00 and 09:00".

[0063] In this embodiment, the collaboration of the rewriting model 220, the intent recognition model 216, and the recommendation model 214 significantly improves the rewriting acceptance rate (the proportion of users actively adopting rewriting suggestions) and the problem resolution rate (enhanced accuracy through rewriting). This collaborative model work in this embodiment balances semantic accuracy and personalization, resulting in a significant reduction in average problem resolution time.

[0064] It should be understood that, in other embodiments of this disclosure, the rewriting model 220 can also rewrite and enhance the user input 210 independently, without the participation of the intent recognition model 216 and the recommendation model 214. It should also be understood that the rewriting model 220 can directly rewrite the user input 210, or it can rewrite the standardized summary extracted based on the user input 210. In other words, when the text summarization model 218 is involved in the processing, the rewriting model 220 can rewrite and semantically enhance the standardized summary extracted by the text summarization model 218, addressing the shortcomings of traditional summarization models in terms of contextual coherence.

[0065] It should also be understood that, in the case where a user selects a recommended query from recommendation model 214, rewriting model 220 can also rewrite the selected recommended query, for example, by adding contextual information.

[0066] In some embodiments, the rewriting model 220 may be a language model. For example, the rewriting model 220 may be the same model as the language model 230, or it may be a different model. For instance, the rewriting model 220 may be a language model with a relatively smaller number of parameters than the language model 230.

[0067] Continue to refer to Figure 2 The dual-path retrieval model 222 can retrieve at least one text fragment associated with the query indicated by the user input from the knowledge base 224, based on user input 210 or enhanced user input generated by the rewriting model 220. This at least one text fragment and the user input can be provided to the language model 230. The language model 230 generates a response 234 to the user input, such as maintenance suggestions, based on the user input and the at least one text fragment.

[0068] In some embodiments, the dual-path retrieval model 222 can retrieve at least one text fragment from the knowledge base 224 based on keyword retrieval and vector retrieval. Keyword retrieval can accurately match hard keywords such as error codes and API names (e.g., "HTTP 502"). Vector retrieval can generalize the query intent and avoid deviation from the query intent. The dual-path retrieval model 222 may also include a weight fusion model, which can perform dynamic weight fusion on the results of keyword retrieval and vector retrieval. For example, the dual-path retrieval model 222 automatically adjusts the score ratio of the results in keyword retrieval and vector retrieval according to the query type. The system according to this embodiment adopts a dual-path retrieval model, which significantly increases the average recall rate and can cover multi-source data such as operation and maintenance manuals and work order systems.

[0069] In some embodiments, keyword retrieval can be implemented based on the BM25 (Best Matching 25) model. Vector retrieval can be implemented based on a fine-tuned Contriever model. The training of the dual-path retrieval model 222 can be accomplished as follows: For keyword retrieval models such as BM25, an operations and maintenance domain keyword dictionary can be constructed, and the keyword retrieval model can be trained based on the operations and maintenance domain keyword dictionary. For vector retrieval models, the model can be fine-tuned using operations and maintenance QA (question answering) based on the Contriever model. For weighted fusion models, the query type can be automatically determined and the retrieval weight score can be adjusted by training a classifier. For example, when the query type tends to be keywords (e.g., user input 210 includes keywords from the keyword dictionary), a higher weight score can be assigned to the text in the keyword retrieval results. Conversely, a higher weight score can be assigned to the text in the vector retrieval results.

[0070] Continue to refer to Figure 2The dual-path retrieval model 222 can provide at least one text fragment obtained from the knowledge base 224 to the fine-ranking model 226, the rearrangement model 228, or the language model 230.

[0071] The fine-ranking model 220 can determine the correlation between each text segment in the at least one text segment obtained by the dual-path retrieval model 222 and the user input 210, and select at least one target text segment from the at least one text segment based on the correlation between each text segment in the at least one text segment and the user input 220. In some embodiments, the fine-ranking model 220 can rank each text segment based on the correlation between each text segment in the at least one text segment obtained by the dual-path retrieval model 222 and the user input 210, and select at least one target text segment from the at least one text segment based on the ranking result. As an example, the dual-path retrieval model 222 obtains 10 text segments, and the fine-ranking model 226 selects 5 text segments from the 10 text segments.

[0072] The fine-running model 220 can provide at least one target text fragment to the re-running model 228 or the language model 230.

[0073] Continue to refer to Figure 2 The reordering model 228 can be configured to determine the weights of each text segment in at least one text segment provided by the dual-path retrieval model 222 or at least one target text segment provided by the fine-ranking model 226. In some embodiments, the reordering model 228 can determine the weights of each text segment based on the relevance between each text segment and the domain corresponding to the query, as well as the timeliness of each text segment.

[0074] In some embodiments, the reordering model 228 can be built based on the ColBERT (Contextualized Late Interaction over BERT) model. For example, a domain-specific question-and-answer pair dataset can be constructed and labeled with relevance scores, thereby allowing the ColBERT model to be trained on the operation and maintenance dataset for domain adaptation to obtain the reordering model 228. That is, using a model fine-tuned based on the ColBERT architecture, the semantic associations of operation and maintenance domain-specific question-and-answer pairs can be learned (e.g., establishing a strong association between "certificate chain verification failed" and "OpenSSL version compatibility document"). The fine-tuned ColBERT model can score the domain relevance of the retrieval results (various text fragments provided to the reordering model 228) to the query indicated by the user input 210 (i.e., the current question), with highly adaptable content (such as error type, service component, or technology stack matching content) receiving higher weights.

[0075] In some embodiments, the reordering model 228 may also include a timeliness scoring module. The timeliness scoring module may embed a time sensitivity analysis component or a time sensitivity classifier to automatically identify and de-weight outdated content. For example, the reordering model 228 may use the timeliness scoring module to lower the priority of documents mentioning deprecated Kubernetes API versions (such as `extensions / v1beta1`). Alternatively, the reordering model 228 may use the timeliness scoring module to degrade the score of solutions containing outdated tool version numbers (such as "Ansible 2.8"). Or, the reordering model 228 may use the timeliness scoring module in conjunction with version metadata from the knowledge base (such as Git commit time and last modified date) to dynamically calculate a content freshness score.

[0076] In some embodiments, the rearrangement model 228 can score each text fragment based on the relevance between each text fragment and the domain corresponding to the query, as well as the timeliness of each text fragment, with higher-scoring text fragments having higher weights.

[0077] In some embodiments, the score for each text fragment is calculated as follows: α × Domain Adaptability Score + β × Timeliness Score + γ × Other Feature Score. The Domain Adaptability Score is determined based on the semantic matching degree (0-1) generated by the ColBERT fine-tuned model, and the Timeliness Score is determined based on a decay function (e.g., `1 / (1+Δt)`) of the content update time and the current time difference. Parameters α, β, and γ are obtained through training with labeled data from operational scenarios. This mechanism ensures that two types of content receive high scores: first, highly adaptable and timely text fragments, such as "Using Calico v3.24 to resolve K8s network policy conflicts"; second, highly adaptable but moderately time-sensitive text fragments, such as general solutions to classic problems (version independent). The following content is downgraded: low-adaptability content, such as domain-independent general technical documents; and low-timeliness content, such as solutions containing deprecated commands (e.g., `kubectl run--generator` is invalid in versions 1.18 and above).

[0078] The information processing system 200 of this disclosure can significantly improve the recall rate of long-tail problems and significantly reduce the number of results that users need to filter again by using the rearrangement model 228.

[0079] The rearrangement model 228 can provide each text segment and its weight (or score) to the language model 230.

[0080] Continue to refer to Figure 2 The language model 230 can generate a response 234 to the user input 210 based on the user input 210, at least one text segment or at least one target text segment, and the weights in each text segment.

[0081] It should be understood that in the system 200 of this disclosure embodiment, one or more of the rejection model 212, recommendation model 214, intent recognition model 216, text summarization model 218, rewriting model 220, dual-path retrieval model 222, fine ranking model 226, and re-ranking model 228 can be models with a small number of parameters pre-trained for the operation and maintenance domain, while the language model 230 can be a model with a large number of parameters pre-trained for the operation and maintenance domain. The system 200 of this disclosure embodiment achieves end-to-end optimization of operation and maintenance vertical question answering through the collaboration of large and small models, and has the following advantages: First, it improves efficiency, for example, the number of tokens for long text processing is significantly reduced, and the inference efficiency is significantly improved; Second, it achieves a breakthrough in accuracy, for example, the key information extraction achieves a high accuracy rate, and the recommendation hit rate is greatly improved; Third, it can realize an automated closed loop, from problem diagnosis (log analysis) to solution recommendation (tool invocation) full-link intelligence, and the proportion of manual intervention is greatly reduced.

[0082] Figure 3 An example architecture for constructing a knowledge base example process 300 according to some embodiments of this disclosure is shown.

[0083] like Figure 3 As shown, in this embodiment of the disclosure, a knowledge base 224 is constructed by dividing the operation and maintenance vertical knowledge document 310 into blocks using a knowledge base block model 320. The operation and maintenance vertical knowledge document 310 may include API manuals, configuration templates, agent usage manuals, and other documents related to the operation and maintenance field. In some embodiments, the knowledge base block model 320 can execute a dynamic block strategy. For example, for technical documents, semantic segmentation is performed using a language model, paragraph-level intent recognition is performed on the technical documents, and semantically coherent text blocks (Chunks) are generated. For example, a document about K8s network configuration can be divided into modules such as {"Network Policy Configuration", "Ingress Controller", "CNI Plugin Selection"}. For example, for structured content such as code snippets and configuration templates, they can be segmented according to predefined rules (such as by character length, delimiter `---`). For example, the Ansible Playbook (a learning manual for the automated operation and maintenance tool Ansible) can be divided into multiple independent task blocks by the `- name:` ​​tag. In some implementations, the knowledge base segmentation model 320 can be trained offline using historical operation and maintenance documents (such as API manuals and configuration templates) to learn how to perform semantic and rule-based segmentation of heterogeneous content. It should be understood that the knowledge base segmentation model 320 can also segment and classify operation and maintenance knowledge documents based on other methods.

[0084] In some embodiments, when new operation and maintenance documents are added to the database, the system automatically triggers a chunking model for processing, generating text blocks with multi-level headings (e.g., "K8s Network Configuration → Calico Plugin → YAML Example"). That is, the chunking of knowledge documents is pre-executed, rather than being performed in real-time when a user queries. Reference Figure 2 Upon receiving user input 210, system 200 directly uses the pre-segmented knowledge base 224 for retrieval without performing real-time segmentation of the documents in knowledge base 224. As an example, for knowledge base documents related to OpenStack, the pre-segmentation results are: [Chunk1] Title: Neutron Infrastructure, Content: Neutron Component Architecture Diagram, Core API List...; [Chunk2] Title: Security Group Configuration, Content: Security Group Rule Syntax Example, Common Configuration Errors...; [Chunk3] Title: VLAN Mode Deployment, Content: vlan_network_type Configuration Steps, Switch Integration Instructions. When a user queries "Security group rules cause SSH connection failure," the system directly retrieves the content of the pre-segmented Chunk2, without needing to parse the PDF document in real time. The knowledge base construction method of this embodiment uses a pre-segmentation mechanism to avoid processing the original document for each query, reducing retrieval latency to the millisecond level.

[0085] like Figure 3 As shown, in some embodiments, the update module 330 updates the documents in the knowledge base 224. In some embodiments, the update module 330 can automatically identify expired segments using a version comparison tool (such as Git Diff) and trigger incremental updates, that is, only incrementally chunking the changed parts without requiring a full reprocessing. This allows for automatic tracking of the version history of the knowledge base content by integrating tools such as Git, ensuring that the chunking results are synchronized with the latest documents. The knowledge base update mechanism of this disclosure ensures strict synchronization between the chunking results and documents through version control, avoiding dirty data such as "documents have been updated but chunks are not synchronized." Furthermore, due to support for automated incremental updates, new documents can be updated at a rate of seconds.

[0086] Figure 4 A flowchart of a process 400 for information processing according to some embodiments of the present disclosure. The process 400 for information processing according to embodiments of the present disclosure, the process 400 in... Figure 1 It is executed at either the client device 110 or the server 120. In this embodiment of the disclosure, implementation at the client device 110 is used as an example for explanation.

[0087] At box 410, client device 110 responds to receiving user input by determining the query indicated by the user input.

[0088] At box 420, client device 110 retrieves at least one text fragment associated with the query from the knowledge base.

[0089] At box 430, client device 110 determines the weight of each text fragment in at least one text fragment based on the relevance between each text fragment in at least one text fragment and the domain corresponding to the query, as well as the timeliness of each text fragment.

[0090] At box 440, client device 110 uses a first machine learning model to generate a response to user input based on user input, at least one text fragment, and the weights of each text fragment.

[0091] In some embodiments of this disclosure, retrieving at least one text fragment associated with a query from a knowledge base includes: in response to receiving user input, determining whether the user input is associated with a predetermined domain; and in response to determining that the user input is associated with a predetermined domain, retrieving at least one text fragment associated with the query from the knowledge base, wherein process 400 further includes: in response to determining that the user input is not associated with a predetermined domain, providing a prompt message to prompt the user that the user input is not associated with a predetermined domain.

[0092] In some embodiments of this disclosure, determining whether a user input is associated with a predetermined domain includes: based on the user input, using a second machine learning model (e.g., a rejection model) to determine the association result between the user input and the predetermined domain and the corresponding confidence score; and in response to the association result indicating that the user input is associated with the predetermined domain and the confidence score being greater than a first threshold, determining that the user input is associated with the predetermined domain; or in response to the association result indicating that the user input is not associated with the predetermined domain and the confidence score being greater than the first threshold, determining that the user input is not associated with the predetermined domain.

[0093] In some embodiments of this disclosure, determining whether a user input is associated with a domain further includes: in response to a confidence score greater than a second threshold and less than a first threshold, extracting key information from the user input using a third machine learning model (e.g., a text summarization model); and based on the key information, re-determining the association result between the user input and a predetermined domain and the corresponding confidence score using a second machine learning model.

[0094] In some embodiments of this disclosure, process 400 further includes: in response to receiving user input, obtaining at least one of the following associated with the user input: user historical behavior data, context information; and based on the user input and at least one of the user historical behavior data and context information, using a fourth machine learning model (e.g., a recommendation model) to determine a set of recommended queries for the user input to choose from. Determining the query indicated by the user input includes determining the query selected by the user from the set of recommended queries.

[0095] In some embodiments of this disclosure, process 400 further includes: generating enhanced user input using a fifth machine learning model (e.g., a language model) based on user input and contextual information related to the user input. Generating a response to the user input includes: generating a response to the user input based on the enhanced user input.

[0096] In some embodiments of this disclosure, generating enhanced user input using a fifth machine learning model includes: determining at least one processing intent indicated by the user input; generating at least one candidate user input using the fifth machine learning model based on the user input and the at least one processing intent; selecting a candidate user input from the at least one candidate user input using a fourth machine learning model; and generating enhanced user input using the fifth machine learning model based on the selected candidate user input and contextual information related to the user input.

[0097] In some embodiments of this disclosure, process 400 further includes: determining whether log information exists in the user input; in response to determining that log information exists in the user input, extracting structured information from the log information using a third machine learning model; and generating a standardized summary for the user input based on the structured information, wherein the acquisition of at least one text fragment is based on the standardized summary, and the generation of a response to the user input is based on the standardized summary.

[0098] In some embodiments of this disclosure, process 400 further includes: determining the correlation between each text segment in at least one text segment and the user input using a sixth machine learning model (e.g., a ranking model) based on user input; and selecting at least one target text segment from the at least one text segment based on the correlation between each text segment in at least one text segment and the user input. Generating a response to the user input includes: generating a response to the user input using a first machine learning model (e.g., a language model) based on the user input, at least one target text segment, and the weights of each text segment in the at least one target text segment.

[0099] In some embodiments of this disclosure, process 400 further includes: determining the processing intent corresponding to the query based on user input, wherein obtaining at least one text fragment associated with the query from the knowledge base includes: obtaining at least one text fragment associated with the query from the knowledge base based on the processing intent.

[0100] The information processing method disclosed in this embodiment has the following beneficial effects: it constructs a large and small model fusion technology system for operation and maintenance verticals, which not only retains the semantic understanding advantages of language models (such as LLM), but also solves the efficiency and professionalism problems of applying large models in operation and maintenance scenarios through domain-specific small models. Furthermore, compared with traditional solutions, log processing speed and recommendation accuracy are significantly improved, and the average fault repair time is significantly reduced.

[0101] Figure 5 A schematic structural block diagram of an apparatus 500 for information processing according to certain embodiments of the present disclosure is shown.

[0102] like Figure 5 As shown, the apparatus 500 includes a query determination module 510, configured to determine a query indicated by user input in response to receiving user input. A text acquisition module 520 is configured to acquire at least one text fragment associated with the query from a knowledge base. A weight determination module 530 is configured to determine the weight of each text fragment in the at least one text fragment based on the relevance between each text fragment and the domain corresponding to the query, and the timeliness of each text fragment. A response generation module 540 is configured to generate a response to the user input using a first machine learning model, based on the user input, the at least one text fragment, and the weights of each text fragment.

[0103] In some embodiments of this disclosure, the apparatus 500 further includes a rejection module configured to determine whether user input is associated with a predetermined domain in response to receiving user input; and a text acquisition module 520 configured to acquire at least one text fragment associated with the query from a knowledge base in response to determining that user input is associated with a predetermined domain, wherein the rejection module is further configured to provide a prompt message to the user in response to determining that user input is not associated with a predetermined domain.

[0104] In some embodiments of this disclosure, the rejection module is configured to: determine the correlation result between the user input and a predetermined domain and the corresponding confidence score based on the user input using a second machine learning model (e.g., a rejection model); and determine that the user input is associated with the predetermined domain in response to the correlation result indicating that the user input is associated with the predetermined domain and the confidence score is greater than a first threshold; or determine that the user input is not associated with the predetermined domain in response to the correlation result indicating that the user input is not associated with the predetermined domain and the confidence score is greater than a first threshold.

[0105] In some embodiments of this disclosure, the rejection module is further configured to: extract key information from the user input using a third machine learning model (e.g., a text summarization model) in response to a confidence score greater than a second threshold and less than a first threshold; and, based on the key information, redetermine the correlation result between the user input and a predetermined domain and the corresponding confidence score using a second machine learning model.

[0106] In some embodiments of this disclosure, apparatus 500 further includes a recommendation module configured to: in response to receiving user input, acquire at least one of the following associated with the user input: user historical behavior data, context information; and, based on the user input and at least one of the user historical behavior data and context information, utilize a fourth machine learning model (e.g., a recommendation model) to determine a set of recommended queries for the user input to choose from. Furthermore, query determination module 510 is also configured to determine the query selected by the user from the set of recommended queries.

[0107] In some embodiments of this disclosure, apparatus 500 further includes a rewriting module configured to generate enhanced user input using a fifth machine learning model (e.g., a language model) based on user input and contextual information associated with the user input. The response generation module 540 is configured to generate a response to the user input based on the enhanced user input.

[0108] In some embodiments of this disclosure, the rewriting module is configured to: determine at least one processing intent indicated by user input; generate at least one candidate user input using a fifth machine learning model based on the user input and the at least one processing intent; select a candidate user input from the at least one candidate user input using a fourth machine learning model; and generate an enhanced user input using the fifth machine learning model based on the selected candidate user input and contextual information related to the user input.

[0109] In some embodiments of this disclosure, the apparatus 500 further includes a summary generation module configured to: determine whether log information exists in the user input; in response to determining that log information exists in the user input, extract structured information from the log information using a third machine learning model; and generate a standardized summary for the user input based on the structured information, wherein the acquisition of at least one text fragment is based on the standardized summary, and the generation of a response to the user input is based on the standardized summary.

[0110] In some embodiments of this disclosure, the apparatus 500 further includes a fine-ranking module configured to: determine the correlation between each text segment in at least one text segment and the user input using a sixth machine learning model (e.g., a fine-ranking model) based on user input; and select at least one target text segment from the at least one text segment based on the correlation between each text segment in the at least one text segment and the user input. The response generation module 540 is configured to: generate a response to the user input using a first machine learning model (e.g., a language model) based on the user input, at least one target text segment, and the weights of each text segment in the at least one target text segment.

[0111] In some embodiments of this disclosure, the apparatus 500 further includes an intent recognition module configured to: determine the processing intent corresponding to the query based on user input, wherein the text acquisition module is configured to: acquire at least one text fragment associated with the query from a knowledge base based on the processing intent.

[0112] The units and / or modules included in device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0113] It should be understood that one or more steps in the above methods can be performed by suitable electronic devices or combinations of electronic devices. Such electronic devices or combinations of electronic devices may include, for example, […]. Figure 1 The first device in the middle is 110. Figure 6 A block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 6The electronic device 600 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 6 The electronic device 600 shown can be used to achieve Figure 1 Client device 110 or Figure 5 The device 500.

[0114] like Figure 6 As shown, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 600.

[0115] Electronic device 600 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data and accessible within electronic device 600.

[0116] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 6 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0117] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.

[0118] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interfaces (not shown).

[0119] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0120] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0121] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0122] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0124] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for information processing, comprising: in response to receiving a user input, determining a query indicated by the user input; obtaining at least one text snippet associated with the query from a knowledge base; determining a weight of each of the at least one text snippet based on a relevance between a domain corresponding to the query and the each of the at least one text snippet, and a timeliness of the each of the at least one text snippet; and generating a reply to the user input based on the user input, the at least one text snippet, and the weight of the each of the at least one text snippet using a first machine learning model.

2. The method of claim 1, wherein obtaining the at least one text snippet associated with the query from the knowledge base comprises: in response to receiving a user input, determining whether the user input is associated with a predetermined domain; and in response to determining that the user input is associated with the predetermined domain, obtaining the at least one text snippet associated with the query from the knowledge base, wherein the method further comprises, in response to determining that the user input is not associated with the predetermined domain, providing a prompt information to prompt the user that the user input is not associated with the predetermined domain.

3. The method of claim 2, wherein determining whether the user input is associated with the predetermined domain comprises: based on the user input, determining a relevance result and a corresponding confidence score of the user input to the predetermined domain using a second machine learning model; and in response to the relevance result indicating that the user input is associated with the predetermined domain and the confidence score being greater than a first threshold, determining that the user input is associated with the predetermined domain; or in response to the relevance result indicating that the user input is not associated with the predetermined domain and the confidence score being greater than the first threshold, determining that the user input is not associated with the predetermined domain.

4. The method of claim 3, wherein determining whether the user input is associated with the predetermined domain further comprises: in response to the confidence score being greater than a second threshold and less than the first threshold, extracting key information in the user input using a third machine learning model; and based on the key information, re-determining the relevance result and the corresponding confidence score of the user input to the predetermined domain using the second machine learning model.

5. The method of claim 1, further comprising: in response to receiving the user input, obtaining at least one of user historical behavior data and context information associated with the user input; and based on the user input and the at least one of user historical behavior data and context information, determining a set of recommended queries for the user input using a fourth machine learning model for the user to select from, and wherein determining the query indicated by the user input comprises determining a query selected by the user from the set of recommended queries.

6. The method of claim 1, further comprising: ​ ​ ​ ​ ​ generating, based on the enhanced user input and context information related to the user input, a reply to the user input, wherein generating the reply to the user input comprises: generating, based on the enhanced user input, a reply to the user input.

7. The method of claim 6, wherein generating the enhanced user input using the fifth machine learning model comprises: determining at least one processing intent indicated by the user input; generating, based on the user input and the at least one processing intent, at least one candidate user input using the fifth machine learning model; and selecting, using the fourth machine learning model, a candidate user input from the at least one candidate user input; and generating, based on the selected candidate user input and context information related to the user input, an enhanced user input using the fifth machine learning model.

8. The method of claim 1, further comprising: determining whether there is log information in the user input; in response to determining that there is log information in the user input, extracting structured information in the log information using a third machine learning model; and generating a standardized summary for the user input based on the structured information, wherein the obtaining of the at least one text snippet is based on the standardized summary, and the generating of the reply to the user input is based on the standardized summary.

9. The method of claim 1, further comprising: determining, based on the user input, a relevance of each text snippet in the at least one text snippet to the user input using a sixth machine learning model; and selecting, based on the relevance of each text snippet in the at least one text snippet to the user input, at least one target text snippet from the at least one text snippet, wherein generating the reply to the user input comprises: generating, based on the user input, the at least one target text snippet, and a weight of each text snippet in the at least one target text snippet, a reply to the user input using the first machine learning model.

10. The method of claim 1, further comprising: determining, based on the user input, a processing intent corresponding to the query, wherein obtaining the at least one text snippet associated with the query from the knowledge base comprises: obtaining the at least one text snippet associated with the query from the knowledge base based on the processing intent.

11. An apparatus for information processing, comprising: a query determination module configured to determine, in response to receiving a user input, a query indicated by the user input; a text obtaining module configured to obtain at least one text snippet associated with the query from a knowledge base; a weight determination module configured to determine a weight of each text snippet in the at least one text snippet based on a relevance between a domain corresponding to the query and each text snippet in the at least one text snippet, and a timeliness of each text snippet; and a reply generation module configured to generate a reply to the user input based on the user input, the at least one text snippet, and the weight of each text snippet in the at least one text snippet. a reply generation module configured to generate, using the first machine learning model, a reply to the user input based on the user input, the at least one text segment, and the weights of the respective text segments.

12. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method of any of claims 1-10.

13. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method of any of claims 1-10.

14. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any of claims 1-10.